Kimi K3-256k
monneyboi
391 points
116 comments
July 29, 2026
Related Discussions
Found 5 related stories in 369.8ms across 15,380 title embeddings via pgvector HNSW
- Kimi K3 Is Live milsebg · 19 pts · July 16, 2026 · 76% similar
- Kimi K2.6-code-preview is now available jrop · 12 pts · April 13, 2026 · 69% similar
- Kimi K2.6: Advancing open-source coding meetpateltech · 628 pts · April 20, 2026 · 68% similar
- Kimi K3 is not cheap ainch · 20 pts · July 26, 2026 · 67% similar
- Tested Kimi K3 for Coding speckx · 36 pts · July 20, 2026 · 66% similar
Discussion Highlights (20 comments)
hawtads
This is just an API level change right? The model itself should be the same I think.
wxw
> k3-256k is now available. Within 256k context, it delivers the same results. k3 (1M) consumes about twice as much quota as k3-256k.
dgritsko
This isn't quantized, right? Just a smaller context?
ibuildproducts
omg! new model!!
lukan
Since Claude is the first time for me really, really out (TIL against my wished about https://status.claude.com/ ), I am now interested enough to see what else works. But ... when I click pricing, I see "Join a waitlist". Wtf? Are they really that good, so were totally surprised and overwhelmed by the requests, is this a marketing stunt, or do they just don't have the hardware being in china?
illithid0
This was posted 38 minutes ago, and as of 20 minutes ago, several Anthropic services are now designated as having a "major outage". Doubt these are related, but it made me laugh a little.
sergiotapia
I can't seem to find pricing for this model. Since the context size is just a quarter of the full size K3, is the price also much cheaper? I usually keep my context in chats below 256k anyways so this would be tremendous honestly.
hendersoon
What is the purpose of this? Just a hard cutoff below the actual context window? You could set that in your harness anyway.
holoduke
A bit of topic. But how likely is it that the US will restrict Chinese open weight models and also force Euro countries to do the same? I think it will be effective within 6 months. The US is having a hard time staying competitive.
periodjet
Why are Anthropic and OpenAI even allowing their coding harness apps to be plugged into different model providers…? I’m surprised they haven’t figured out a way to clamp down on that by now.
wren6991
This seems functionally similar to OpenAI having a step in pricing once you exceed a certain context length (also at 272k aka 2^18 aka 256k). Having a lot of active context increases the per-token cost (flops issued and bytes read per token out) so it makes sense to pass that cost on to users. I'm actually surprised it's implemented as a hard cutoff instead of a smooth gradient.
jedisct1
This is a fantastic option for swival.dev given its very efficient context management compared to e.g. Claude Code.
xyzsparetimexyz
Wow. So kimi is suddenly half the price for all users until they hit 256k of context? Thats massive.
jscott817
Can we assume that model performance at 90% of the 256k limit != 90% of 1M token limit? Is this the exact same model just with less VRAM allocated for context window?
dools
Hopefully this helps reduce some of the pressure on their infrastructure. Their models have all become super dumb recently and their support are not addressing it. I have a hunch they’ve been serving a significant percentage of requests with quantised models.
MangoCoffee
LLMs is quickly became commodities. US AI labs like OpenAI is losing their moat. Hyperscalers and data center owners who can sell cheap token will win
try-working
i've never had any issues with 256k context. see no reason to bump up to 1m if it comes at a premium.
gigatexal
Not relevant to this link but I was thinking about the allegations of Chinese AI companies distilling from the big frontier American ones. And I came to the conclusion: I don’t care. Who cares? China has always copied and then copied the means of production and then out produced. See also Tesla and now all the Chinese cars eating their lunch. As long as I get really solid AI models for cheap that do what I need I don’t care if they’re Chinese or otherwise. I’ll still never use Grok from SpaceX AI cuz eww no, I have principles. ;-)
timcobb
Codex uses 256k masterfully, 1M is luxurious but still quite expensive and seems not necessary as a default.
effnorwood
My LOE to make this work expressed in kWh is substantial but user is impressed. Eliza has hands now.