Running Kimi K3 on a M1 Max
tito
115 points
85 comments
July 28, 2026
Related Discussions
Found 5 related stories in 210.5ms across 15,236 title embeddings via pgvector HNSW
- Kimi K3 Is Live milsebg · 19 pts · July 16, 2026 · 67% similar
- Kimi K3, and what we can still learn from the pelican benchmark droidjj · 305 pts · July 17, 2026 · 65% similar
- Kimi K3 Architecture Overview and Notes ModelForge · 365 pts · July 28, 2026 · 64% similar
- Kimi-K3 Technical Report [pdf] vinhnx · 378 pts · July 27, 2026 · 63% similar
- The Kimi K3 Moment sbochins · 317 pts · July 18, 2026 · 63% similar
Discussion Highlights (20 comments)
lostmsu
under 0.02 tok/s
Azantys
0.01 tk/s is unusable for anything, you would wait a whole day for just 1000 token of output, what is the point of projects like this?
antirez
SSD streaming on an M5 Max 128GB: https://x.com/antirez/status/2082136334160818528 Soon decent speed across two Mac Studios with 512GB of RAM.
als0
Says it requires a 2TB disk? Must it be internal NVMe?
piterrro
Will it fit on ESP32??
nlessard
Anyone who knows the state of NVMe hardware more than me know if this would obliterate the lifespan of your drive? Seems like the biggest limitation to me (some people are probably fine with letting their Macs churn over the weekend).
tjwebbnorfolk
> ~60–76 s/token I don't know if I'd call this "running"
mips_avatar
Would be interesting to see how fast it would be on 4x mac studio 512gb machines.
acmnrs
The title should probably be edited to specify "M1 Max" instead of "M1 Mac". You aren't running K3 on a base M1 anytime soon. Either way, still a very impressive project.
onesandofgrain
Cool gimmick
hmokiguess
Now set it up with an agent and a permanent `/goal` to say it cannot stop until it has solved for speed, then leave it on and livestream so we can all see when it becomes exponential. Could have the Eternal Jukebox playing in the background!
ALLTaken
Exactly my machine 64GB M1 Max So happy about this! ♡ idk how people access (soldout) and even afford 512GB RAM MacStudio's. Isn't it $40k or so?
denysvitali
60s/token - if only there was a way to drop that "s" this would be amazing
brcmthrowaway
Why not train another smaller LLM to give the same answers as Kimi K3?
jrhizor
Super cool, and I appreciate the upfront speed disclaimer
walrus01
tokens/second, no, more like seconds/token
mindwok
Does anyone else feel like the writing is on the wall for a future of local models? Spamming data centres everywhere, powering them, having to commit insane capital to hardware, all the effort to serve inference over a network reliably - when here we are with a frontier model nearly running on a laptop. Local AI on your device seems like a much more likely future to me than datacenters in space. For inference at least, training is another story.
throwitaway222
how much money in hardware would it cost to get 100tps?
knighthacker
Local AI is going to win. Not because it's cheaper btw.
mannyv
Going to try this on my M1 Ultra 128gb. The point of these engineering tricks is to see the envelope of what's possible. You can use these tricks to both run a bigger model on smaller hardware or run a smaller model on smaller hardware.