Kimi K3 Architecture Overview and Notes
ModelForge
365 points
62 comments
July 28, 2026
Related Discussions
Found 5 related stories in 337.2ms across 15,236 title embeddings via pgvector HNSW
- Kimi K3, and what we can still learn from the pelican benchmark droidjj · 305 pts · July 17, 2026 · 73% similar
- Kimi-K3 Technical Report [pdf] vinhnx · 378 pts · July 27, 2026 · 72% similar
- Kimi K3 Is Live milsebg · 19 pts · July 16, 2026 · 70% similar
- Kimi K3 Intelligence, Performance and Price Analysis theanonymousone · 51 pts · July 16, 2026 · 69% similar
- The Kimi K3 Moment sbochins · 317 pts · July 18, 2026 · 68% similar
Discussion Highlights (11 comments)
gokohl
Interesting that they went NoPE everywhere — everyone else hedges with RoPE in the local layers. Feels like the linear-attention stuff (Kimi Delta) is quietly doing the positional work so they can get away with it. Curious to see if it holds up at frontier scale.
alealvarezarg
Great breakdown. After using Kimi extensively, it's fascinating to see how architectural choices like KDA and NoPE translate into such strong real-world performance. Really impressive engineering.
thatsgcasey
Sabastian Raschka is one of the great LLM researchers/authors. I highly recommend his substack
souravsspace
i like your detailed breakdown. Thanks. <3
Ilaurens
"Interestingly, Kimi K3 got rid of all RoPE layers and uses NoPE (No Positional Embeddings) everywhere instead." It just baffles me that this even works at all. Doesn't it just become a token soup? Is attention that precise that a second token can tell its the second token just because it learns to accumulate something in the embedding space without any sort of inductive bias?
constantlm
So, unlike what leaders of western labs labs would like you to believe (that Kimi is just the result of distillation attacks), they are introducing new and novel approaches.
RobLach
Concise. Nice.
rglover
Just tried K3 out for the first time today and it's a legitimate threat. Temporarily (maybe permanently) using it as my daily driver but it's wild how comparable it is to Opus 4.7/4.8 (what's been my go to for a bit now—wrote a quick post on what I found today [1]). [1] https://graybearding.bearblog.dev/kimi-k3-is-insane/
mickael-kerjean
Genuine question: how reproducible / usable / verifiable are these architectures from the published documentation? Are they similar to PDF/DWG/PSD specifications, where the format look like an open spec at first sight until you attempt to implement it and realize the crucial implementation details are undocumented?
augment_me
I feel like the Kimi team is amongst the best in the industry to pick and choose what is meaningful from the other models. For example, avoiding the expensive and empirically uncertain mHC in favor of simpler residuals. Latent MoE. My only doubts are around Linear Attention instead of DSA as this is inherently lossy. You are kind of banking on that your query is inherently in the embedding space of the model already and can be lossy.
gboss
Anybody getting the result that Kimi 3 is more expensive than Opus 5 or Sol on Cursor? Pretty sure Kimi 3 sucked up a good chunk of my ultimate plan in a few prompts. Anyone have any tools or ways to understand per model usage towards cursor subscriptions? I know there are alternatives to cursor just haven’t made the move yet. (Edit spelling)