Advancing the price-performance frontier with GPT‑5.6

tedsanders 541 points 352 comments July 30, 2026
openai.com · View on Hacker News

Discussion Highlights (20 comments)

bakugo

> GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less Looks like the Chinese models are really making a dent. Having 3 different price categories with the "most affordable" one still costing more than GLM 5.2 never made sense.

simonw

> The kernel work helped reduce the end-to-end cost of serving the model by 20%, while its experiments increased token-generation efficiency by more than 15%. If the cost of serving GPT-5.6 just dropped by 20%, does that add up to literally billions of dollars in savings per month? We know Anthropic spend $1.25 billion renting inference capacity from SpaceX (in two Colossus datacenters) from the SpaceX IPO, but we don't know how much of Anthropic's inference capacity that is (presumably a small fraction, since they were operating on top of AWS and other providers before the SpaceX deal.) I've not seen any numbers that hint at OpenAI's per-month inference bill, but surely that has to be in the multiple billions of dollars as well. So 20% is a really, really big deal.

pavpanchekha

Making Luna, which was already very cheap and extremely capable, 5x cheaper is crazy. I use Sol at work but Luna at home, and while there's definitely a difference, it doesn't feel like night-and-day. After a year of ever-increasing prices it suddenly feels (between this, Kimi K3, GLM 5.2) that prices are falling again.

measurablefunc

Model segmentation & distillation like this that asks the consumers to pick exactly which version of the algorithm will solve their problem is evidence for lack of intelligence instead of its presence.

sidcool

They don't mention Grok at all.

goldsmith112

Not sure who would use Terra anymore. Pair Luna High/Xhigh with Sol Medium and that's your power stack

swingboy

This is awesome. Luna is a pretty great model on xhigh.

preommr

> Starting today, GPT‑5.6 Luna, our fastest and most affordable model, will cost 80% less, I don't have the words. I genuinely thought we were in a stage where we were plateauing and going in for 5-10% improvements over months. Seeing spikes like this makes me question about where the floor really is.

gentlewater

This is awesome. I’ve recently set up my opencode to use 5.6 terra for my main agent, who delegates work to a 5.6 Luna coder agent. So far it seems to work well, and reduce costs a lot. With this price reduction, it will work a whole lot better. Perhaps I can get my github copilot quota to last the whole month now.

tosh

80% price cut for luna is a very aggressive pricing move makes it by far the best choice for most workloads that do not need bleeding edge intelligence (reminder: luna can be comparable to opus 5!)

firasd

This is one of the things OpenAI has been focused on for an year or so that led to the doomed autoswitcher in ChatGPT .com (switching models based on estimated task complexity) that was quickly reverted Whereas Google with Gemini 3.x, Anthropic with Fable etc are happy to just go for 'big model with dense params' It's hard to guess from the outside of course but just this kind of talking points focus on GPU efficacy is what we see from OpenAI and Chinese open source labs more often than from Anthropic or Google Deepmind and this benchmark chart seems to concur

bob1029

This feels like the dialup->broadband transition to me. I was already a huge proponent of Luna for things like deep research. Being able to run 5x more for the same cost is simply bananas. We are already running 10 parallel agents for hypothesis generation. I cannot imagine 50. The statistics become much more interesting & powerful when you can run so many samples of the exact same prompt+model without breaking the bank.

alvis

Basically lunar at extra level can cover all use cases scenarios other than those requiring opus up. Goodbye sonnet and haiku

sosodev

Looks like I might have a reason to use something other than Deepseek V4 Flash.

GodelNumbering

"Half the money I spend on advertising is wasted; the trouble is I don't know which half." -John Wanamaker This applies even more strongly to model choosing. I know for a fact that majority of my work doesn't require a very strong model, but separating the trivial and non-trivial tasks is a famously hard problem (if at all decidable).

quirino

I generally just check the Price/Performance graph on Openrouter: https://openrouter.ai/rankings#performance#benchmarks . Activate the "Show Pareto" toggle on the right. I was still using GLM-5.2 in my personal projects, but this just made Luna a very easy choice.

wronex

What are your use case for these? I’m manly interested in coding where more capability is better - give me a 10x model at 10x the price and I’ll take it. A worse model at very low cost has no appeal to me. At-least not for coding. Translation maybe? OCR?

kingstnap

Those prices on luna are killer. Haiku was already in a ditch. But this is coming straight for the jugular of a ton of models on openrouter.

fractorial

It would appear that rolling my own Anthropic-free harness / serving stack with a closed-weight carve out for Codex models is an absolute win.

peheje

Might just resub. Will experiment with Luna next sessions. 5 h window is not working very well for me. But if I can drop down to Luna at 20-30 % left and comfortably ride out the wave then.. that might just work.

Semantic search powered by Rivestack pgvector
15,510 stories · 144,699 chunks indexed