SpaceXAI launches Grok 4.7: larger base model at same pricing as 4.6

SpaceXAI has released Grok 4.7, its new flagship model for coding, agentic tasks and knowledge work. The core change is a larger base model and a longer reinforcement-learning run on difficult tasks, while pricing stays unchanged at $2 per million input tokens and $6 per million output tokens (approximately 7.4 and 22 shekels respectively). The model is available now through the xAI API, Cursor, Grok Build, OpenRouter, Vercel and Cloudflare.
According to the developer documentation, four main changes separate 4.7 from its predecessor: an entirely new base model rather than a reuse of the 4.6 base, an extended RL run on tasks that take hours to complete, improved self-verification and long-context handling, and dedicated training for the Grok Bot harness for conversation and knowledge work. The context window stands at 500 thousand tokens, and the API offers four reasoning-effort levels up to xHigh.
In the launch benchmark table Grok 4.7 improves on 4.6 across every row. The largest jump appears on Terminal-Bench 4.0, rising from 20.3% to 38.0%. On EEBench it reaches 64.0%, the highest score in the table. On Harvey, a legal-agent benchmark, it scores 19.6% versus 6.7% for Fable 5.1 Max. The lead is not absolute, however: Fable 5.1 Max leads on four of seven tests, including 57.9% on Terminal-Bench, and GPT-5.6 Sol Max holds the best DeepSWE v1.1 result at 72.7%. All scores are vendor-reported, meaning they were published by the company itself without external verification.
The price axis is where the primary competitive advantage sits. Fable 5.1 Max costs 5 times as much on input and roughly 8.3 times as much on output. GPT-5.6 Sol Max costs 2 times as much on input and about 3.3 times as much on output. On the CursorBench 4.0 cost-per-task chart, SpaceXAI positions Grok 4.7 at the front of the price-performance curve. On GDPval, which measures professional knowledge work, Grok 4.7 xHigh scores 1,695 Elo, up from 1,605 for 4.6 High. Fable 5.1 Max leads at 1,735, while GPT-6 Astra max sits at 1,542. The company also notes improvements in document and presentation generation.
Security arrives with an entirely new safeguard stack. SpaceXAI describes Grok 4.7 as the strongest model it has tested on refusals and jailbreak resistance. On LatchBio biosafety it scores 62.4%. On HackerBench v0.3, an internal benchmark for dangerous cyber tasks, the model allowed 3.3% of problematic dual-use prompts to pass; the company claims it barely blocks legitimate security work. Selected cyber partners receive invite-only access to red-team capabilities for defensive research.
On pricing and availability, Grok 4.7 Fast is the same model running on faster infrastructure, delivering twice the output speed at double the price. It runs only in Cursor and Grok Build, not on the public API, and is excluded from the Grok Build free tier. A US regional endpoint keeps inference within US borders at a 10% premium. SpaceXAI recommends setting a prompt_cache_key for reliable cache hits.