GPT-5.6 vs. Claude Fable 5 for Physical AI, which performs best?

mbauman 90 points 19 comments July 29, 2026
juliahub.com · View on Hacker News

Discussion Highlights (11 comments)

jespinel

Nice! It is missing Codex in the agent harnesses comparison IMO.

grim_io

I'd expect google to do well here, since they were historically strong at multimodal and physics.

gizmodo59

Yet another "benchmark to promote their own harness"

giwook

Please forgive my naivety, but are world models (once they are in a consumer-ready form) expected to outperform any currently existing LLM on these sorts of tasks (i.e. of the physical world)?

hartator

It's kind of interesting this is already out of data as it's missing Kimi 3 and Opus 5.

xnorswap

A really frustrating partial presentation, given an apparent lack of testing with a spread of efforts for each model. Given that there's no reason to believe that Fable's xhigh is comparable to GPT-sol's xhigh, or Opus xhigh, for that matter, it would be far more useful to see the effort level where these tasks no longer achieved their goals.

arisAlexis

Google with apptronic should have good models soon

effnorwood

Define "best" and "performs"

draginol

So Fable "won" but it cost $124.76 for marginal performance benefits over the $22.56 5.6 Sol run.

DwarvenEngineer

Honestly, I just hate the term "physical AI". They're robots. It's unfortunate that we had to adopt a term with the words AI in it, just to get investor's attention.

busssard

compare to opus-5 ?

Semantic search powered by Rivestack pgvector
15,380 stories · 143,452 chunks indexed