Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA

piotrgrabowski 478 points 275 comments July 21, 2026
fireworks.ai · View on Hacker News

Discussion Highlights (20 comments)

Sphax

what are my options if i want to use a router like this ? who provides one ?

culi

A third the cost, open source, and won't refuse every other request because of some vague possible connection to cybersecurity concerns.

guessmyname

Why SoTA (uppercase “T”) instead of SotA (lowercase “T”) ? “State of [T]he Art” versus “State of [t]he Art”. If not SotA then at least SOTA, which is more accurate.

stingraycharles

As always, benchmarks rarely paint the whole picture. It also seems like this article is somewhat biased, eg when Fable and Kimi are close but Fable wins it’s “dead heat”, but when Kimi wins it’s “Kimi wins”. GPT 5.6 seems to be missing as well. I am really eager to give Kimi K3 a try, but I’ll reserve my judgement until I’ve worked with it for at least a few days.

OutOfHere

They forgot to compare and incorporate GPT-5.6-Sol.

jrflo

Hmmm, a company that hosts open models is telling us how good open models are...

hawtads

For model routers, do they have to retrain the routing model every time a new LLM is released?

mattvr

Anyone have routing harnesses like this describes with Claude Code? Or other good routing platform recommendations? (yes, I know this article is about an oracle router)

felipeerias

Mythos/Fable was the state of the art back in March, if not earlier.

refulgentis

AI written drivel by an inference provider, hyping an open model that isn’t open, otherwise the inference provider would be hosting it, and the inference provider also does not have this router they trumpet available, in any form. I’d be amenable to it but I’m already aware of them lying about token throughout by 50-100% on AA and on Twitter, even when you get a dedicated server. Just too much lying piled up to extend courtesy here, even though you’re my favorite open model provider. Please, get your house in order, and if the house is so big that marketing drove this set, ask them to slow down a bit and market things that people can actually use / you actually provide.

nharada

Is there something specifically with Kimi that's better here? As far as I know Kimi pricing is about the same as Sonnet 5 -- what happens if you use that model and Fable instead? Or Grok 4.5 which is even cheaper?

arjie

Interesting. So the latest in the technology now is this model routing thing. Cursor estimated Composer + Fable works much better than Fable alone. And here K3 + Fable is supposedly better. Interesting.

lvl155

It is not SOTA. Give me a break. Sure, run it on Cerebras to get speed but that’s pretty much its advantage.

apatheticonion

I love the Chinese models. I use DeepSeek exclusively and now Kimi K3 offers a great planning assistant for more advanced coding tasks. DeepSeek v4 Flash is extremely fast and is able to handle pretty much anything I've thrown at it (I use mostly Rust, PSQL, Angular and Terraform). I self host Bifrost as my LLM gateway, though I wish LLM vendors would do monthly/daily automatic billing (like VPS providers do) rather than prepaid + auto-top up. It's annoying maintaining a non-refundable minimum balance across vendors, I would rather be billed for my exact usage. OpenRouter helps, but I don't really like it as a service and not a fan of the mark up.

JSR_FDED

Very interesting. They test Kimi K3 and Fable on a set of approx 1000 tasks grouped into 5 areas (SWE, Legal, etc). They put a router model in front that predicts whether Kimi or Fable is going to give a better cost for a correct result. (They believe that ultimately such a router model should be continuously trained on your own workloads so it makes the best decisions for you). Their router chose Kimi the majority of the time (72% in one category, all the way to 96% in another category), leading to cost savings in every category (from 1.5x to 50x depending).

sbinnee

Openrouter also features routing. Routing is indeed an option if you don’t need consistent behavior and allow switching models.

luciana1u

we put a router model in front of two other models so the router can decide which model is better at deciding things. next we'll need a router for the router and eventually the entire internet is just routers routing routers to other routers

Buttons840

I have several thoughts about this, which I'll just iterate: 1) US export bans have made it so that Chinese companies have to compete using less-than-state-of-the-art hardware. This has forced Chinese companies to build more cost efficient models. Whereas, US companies have moreso tried to be state of the art by spending more money than anyone else on state-of-the-art hardware. 2) Xi Jinping has called for more open AI models (not to be confused with the closed models of OpenAI), and I'm happy to see a powerful world leader advocating for open-weight AI models. Whereas, the US seems likely to just ban models. 3) My impression is that, if China surpasses the US in AI development, there will basically be nothing that the US does better than the rest of the world--except for military spending--we spend a lot, but we don't necessarily spend well (something something Iran). I mean, if the US is no longer a tech leader in the world, like... what are we a leader at? Manufacturing? Healthcare? LOL. Are we a leader in any industry or by any metric? I wonder if China is attempting to remove the last jewel in the USA's crown with these AI releases. 4) It must be refreshing for companies to have access to a new model that isn't going to get pulled because the government bans it 2 days after release. And it's open-weight so it wont go away--amazing--what a shift in the market. 5) If it becomes clear that open-weight models are the future of AI, will that pop a huge bubble in the US economy? Maybe. But, on the other hand, these companies aren't just training AIs, they are also building data centers which will remain valuable no matter what happens.

johnhess

Was this an out of sample test of the router or was it trained on these specific use cases/eval suites?

skybrian

The article is about the best you could theoretically do with a perfect router. The takeaway is that trying to build a good router is worth doing. But it's unlikely to be a perfect router.

Semantic search powered by Rivestack pgvector
14,369 stories · 134,336 chunks indexed