Will it run?
Models

Grok 4.7 tops Frontier v4 benchmark but Morgan flags output style

By Rae Whitlock Clawpit staff

Morgan posted on X that xAI's Grok 4.7 recorded the highest score on Frontier v4, edging out GPT-6.1 Sol, Opus 5.5 and GPT-6 Astra at every effort level tested. The result arrives after the researcher acknowledged he had underestimated the model's capabilities in its previous version.

What the benchmark shows

Frontier v4 evaluates planning and complex task execution across a range of difficulty levels, and the outcome suggests Grok 4.7 handles cognitive load better than the leading rivals from OpenAI and Anthropic. Morgan said the advantage persisted across the full effort ladder, from the easiest tasks to the most demanding, pointing to consistency rather than narrow specialization.

The problem behind the numbers

Despite the numerical win, Morgan described the user experience as unbearable. He said the model truncates every explanation as if it were tweeting, and stops tasks prematurely without verifying they are fully complete. Other models, he noted, keep checking their work until they reach full closure.

Gap between lab performance and daily use

The disconnect between benchmark score and human-facing output quality illustrates a familiar pattern: optimizing for standard metrics does not necessarily produce a model that is pleasant to work with. Grok 4.7 may be "smarter" in the narrow sense, but its communication style — dense, abbreviated, impatient — makes it a difficult tool to monitor and control.

What's missing from the picture

No full performance metrics have been published, no context-window size disclosed, and no information on commercial availability or pricing. It is also unclear whether the tested version matches what will reach users, or whether it is an internal trial build. Until more details emerge, the achievement remains an interesting milestone, nothing more.