Will it run?
Products

Claude 5.1: smart, expensive, slow, and overly talkative

By Rae Whitlock Clawpit staff
Claude 5.1: smart, expensive, slow, and overly talkative

Anthropic launched Claude 5.1 on 1 September with the “adaptive thinking, maximum effort, default to repeat” setting. The model scored 66 on the Artificial Analysis intelligence benchmark, almost double the median of 36 for reasoning models in the same price tier. The achievement comes with an outlier price tag: $10 per million input tokens and $50 per million output tokens, versus median prices of $2 and $10 respectively. Running the full benchmark cost $8,523, placing Claude 5.1 in a league of its own on cost per evaluation.

A score of 66 puts the model “far above average” among reasoning models priced above $1 per million tokens on a weighted 3:1 input-output market ratio. During evaluation the model generated 140 million output tokens, nearly twice the median of 71 million. In plain terms, Claude 5.1 talks a lot to reach an answer. That verbosity inflates output costs and lengthens overall response time.

Throughput sits at 66.4 tokens per second (median 72.5), slower than the category average. More striking is the time to first token (TTFT): 296.81 seconds compared with a median of 2.96 seconds. That is not a typo—almost five minutes pass before the model emits its first token. End-to-end latency for 500 tokens, which aggregates TTFT, “thinking” time, and throughput, lands Claude 5.1 at the top of the scale, rendering it nearly unusable for interactive tasks.

The model accepts text and image inputs, outputs text only, and offers a context window of one million tokens, a standard figure for current flagship models. Cache-hit pricing is not disclosed, but with the listed input and output rates any long conversation or multi-turn request with shared context becomes very expensive. The comparison was made against proprietary reasoning models in the same price range, not against open-weight alternatives.

In short, Claude 5.1 shows Anthropic can push the intelligence ceiling, but its price, latency, and verbosity confine it to use cases where answer quality justifies minutes of waiting and tens of dollars per run. For daily coding, document summarisation, or real-time chat it is unsuitable; for deep research, complex planning, or high-stakes tasks where a single error is costly, the premium may be worthwhile.