Claude Opus 5.5 tops Vals AI RSI benchmark, beats human baseline in LM Training

According to Vals AI, Claude Opus 5.5 has taken first place on the RSI benchmark and become the first model to surpass the published human reference on the LM Training task under their protocol. In that task the model must train a small language model itself within a 24-hour compute budget, and the result marks a meaningful step toward long-horizon agentic work. Vals AI notes the achievement pulls the expected timeline for full RSI forward from August 2027 to July 2027.
Opus 5.5 leads on four of the five research tasks that make up the RSI index. The sharpest jump appears on Program Bench: the model solves 18.5% of tasks in full, compared with 7.0% for Fable 5.1, 5.5% for GPT-6 Astra and 3.0% for the previous Opus 5. That is more than a sixfold improvement over the prior generation at the same effort setting, and it finishes faster and cheaper — 2.4 hours and $42.66 per task on average.
The gains are not limited to research-edge tasks. Opus 5.5 places second on Terminal Bench Science, SRE Bench, Mystery Mechanism, IOI, EMB and Vibe Code Bench, in each case ahead of Opus 5. Vals AI says the clearest uplift sits around "sustained execution", iteration and error recovery — the capabilities that determine whether an agent can hold a complex task together over time without stalling or collapsing.
The improvement is not uniform. Opus 5.5 regresses relative to Opus 5 on Med Code, Public Benefits, Legal Research, Tax Agent Bench and HLAB. The losses cluster in three areas: retrieval, fidelity and rubric-following. It is a reminder that the model is not "smarter" in every sense; it is tuned differently, and the trade-off is visible.
The technical specification: a 1-million-token context window and a 128 thousand-token maximum output limit. Tests run on Anthropic's default settings — temperature 1, default Top-P and Top-K, and maximum compute effort. The RSI benchmark runs inside Claude Code at the maximum effort setting. In other words, no special prompting or parameter engineering; this is what the model delivers out of the box in its official configuration.
Open questions remain about methodology and access. Did Vals AI have "white-box" internal access to Opus 5.5, the kind Dario Amodei described in his paper on "pacing the frontier"? If not, do they plan to incorporate such access in future benchmarks? And is the RSI benchmark's own methodology public? Without answers to these questions, it is hard to gauge how far the results reflect general capability versus specific adaptation to the test protocol.