Kimi K3 is not cheap

ainch 20 points 22 comments July 26, 2026
www.alexinch.com · View on Hacker News

Discussion Highlights (9 comments)

ronsor

It's cheap because it won't refuse random tasks. You can't get rid of nannying at any price beyond training your own model, and relative to that, K3 is cheap.

cloudie78

It’s cheap.

SwellJoe

In my testing, I'm finding it more expensive than Opus 4.8/5 and GPT 5.6 Sol at API rates, because it chews so much. And, their plan (at least the $19 tier) is much less generous than the ChatGPT $20 plan, like an order of magnitude less, it's basically a demo not a useful amount of usage.

himata4113

it IS cheap (once the weights are released) and it will only become CHEAPER. For around $3700 a month (via loan purchased hardware + energy cost) you can run around 32 concurrent instances of kimi k3 which can generate nearly a 6.9 billion tokens a day. This is napkin math since I'm mostly just extrapolating from glm 5.2 by assuming it's twice as heavy to serve in every single measurement, but I believe you can easily achieve 2500tok/s aggregate compared to 4500tok/s and up to 8000tok/s for glm5.2. with nvidia r100 you are likely going to be able to push that number even higher while the cost of hardware appears to be relatively the same, so far I am seeing 21% premium from supermicro which is twice as fast and has nearly twice the vram.

CamperBob2

A lot hinges on what happens tomorrow. I'll believe they'll open the weights when I see the files appear on HF (and when somebody with 24 RTX6000s or whatever reports that they are indeed as good as the closed version.)

coder543

This article seems premature to post. Right now, the price is arbitrarily set by a single provider. Why wouldn't Moonshot collect extra revenue during this exclusivity period when they knew there would be hype? The model weights are supposed to release tomorrow. Over the next several weeks, I would expect competition among open weight providers to drive down the cost, as I've seen happen with other open weight model releases.

sroerick

It feels like this "Kimi is a token hog" meme is 100% astroturfed by Anthropic. It's cheap. Believe your own eyes.

ofjcihen

I mean you have a chart showing that it’s cheaper than the other models and it also does what I want without argument. Additionally, I fully expect the frontier labs to continue increasing prices to meet the profit margins they need to to continue existing.

jszymborski

K2.6 is cheaper than GLM5.2 (at least on DeepInfra) and I've found it works as good as Sonnet for my purposes. Both tend to think themselves into circles a bit and aren't super token efficient, but I've found GLM5.2 much worse on this count making K2.6 even cheaper than the per token price would make seem.

Semantic search powered by Rivestack pgvector
14,941 stories · 139,598 chunks indexed