Grok 4.7
meetpateltech
536 points
463 comments
September 21, 2026
Related Discussions
Found 5 related stories in 78.6ms across 7,302 title embeddings via pgvector HNSW
- Grok 4.6 iLuddite · 486 pts · August 12, 2026 · 97% similar
- SpaceXAI: Grok 4.6 theanonymousone · 29 pts · August 12, 2026 · 79% similar
- Grok Bot rvz · 230 pts · August 11, 2026 · 79% similar
- Grok 4.7 Intelligence, Performance and Price Analysis theanonymousone · 13 pts · September 21, 2026 · 78% similar
- Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index wertyk · 324 pts · August 12, 2026 · 76% similar
Discussion Highlights (20 comments)
ls1911
after using cursor grok & trae.ai for several months , grok curor is highly superior results to trae.ai
simianwords
I guess it’s only my opinion but having used grok for personal chat: it’s by far the worst one amongst Claude, ChatGPT and even Deepseek, Gemini etc. The personality is bland and it doesn’t work nearly as hard or even tries to help.
moojacob
Apparently Grok 4.7 has 40% more weights than Grok 4.6, but the price ($6 output token, $2 input) is the same. Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise. However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby. My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.
kristofferR
What's with the deceptive graph on top? Not including Astra can't have been an oversight, did the model compare poorly to it?
sidgtm
In my experience Grok especially inside Grok build is pretty solid choice, it’s a no nonsense model and stays on its course. Another surface where I truly enjoy the experience of using Grok model is Grok bot
vessenes
Nice to see this release cadence increasing and some continued improvement in quality. I am guessing these models are basically still outcomes of the cursor team integrating with the massive amount of compute they now own: I’d imagine we will see significant step up improvements with grok 5 later this year as the team gets more experienced and confident with larger training deployments. Here’s hoping for another competitive frontier model!
Tsarp
Waiting on simonw "Generate an SVG of a pelican riding a bicycle " benchmark to judge this model
Saline9515
I tried in Omp (Oh-my-pi), and so far it's really problematic. It will loop in thinking mode ("Let me implement those fixes: Fix 1, Fix 2, Fix 3 .... Fix 80, Fix 81"), ignore the AGENTS.md instructions, corrupt plan files, etc etc... I have 5.6 Sol as advisor/watchdog, and it blocks every turn, I never saw this. Quite a shame, 4.6 wasn't so bad.
andsoitis
Congratulations to the team!
AM1010101
Did 4.6 not have an x-high reasoning level? Why are they comparing 4.7 x-high with 4.6 high?
maz1b
Either way, the fact that xAI or SpaceXAI or whatever the name is, I can commend the team behind it on their rapid ascent and progress by being close and or on the frontier in several respects.
6thbit
( why is the x-axis on the first chart in descending order ? )
meerita
Grok it's really expensive. I'm getting really amazing results using DeepSeek 4.1 Flash for fraction of the price.
zug_zug
Well I "tried it out" I asked it one question, and it gave no answer and said "Sign up to use more!" I don't think I'll be doing that, no. I can't think of a single dimension grok is winning on (capability, cost, voice), but want to stay open-minded -- anybody want to vouch for its capabilities in any domain?
simonw
$2/million inout and $6/million output but I couldn't see any pricing information for cached input tokens?
MuffinFlavored
If the CursorBench 4.0 score diagram is the headline, I read it as "Grok 4.7 xHigh is almost the same as Fable5.1 on low". Is there a metric for like... time taken when comparing these two? I see score and cost. If Fable5.1 can knock it out more quickly on low but Grok4.7 might take twice as long to stumble through a problem (and leave behind a bunch of yucky comments or un-needed extra unit tests), are they really comparable? Or like... the "quality" of the solution? "It works" versus "it's unmaintainable/very messy/hacky".
WarmWash
Good thing they used 5.6 sol instead of Astra for benchmarks, the EEbench one is crazy[1] [1] https://eebench.org/
saejox
Not even close to astra. Astra is something else. It is expensive, but uses way fewer tokens do my tasks. xAI missed its chance, Ball is on Anthropic's court.
simonw
https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - default reasoning level. Here's reasoning level high: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... For some reason reasoning effort low and medium used similar numbers of tokens, and xhigh used less than high. I think I need to try without OpenRouter in the middle. UPDATE: I tried again with the xAI API directly: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - not a great deal of difference between reasoning levels, and this time xhigh and low used the same number of reasoning tokens for some reason. For comparison here's a fresh run against Grok 4.6: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
dom96
It’s a shame this model has such negative political baggage associated with it. It’s the only one I decided not to run in my LLM benchmarks[1]. 1 - https://bench.killswitch-lang.org