Will it run?
Models

Artificial Analysis launches new benchmark leaderboard for Wisedocs MLCR

By Rae Whitlock Clawpit staff
Artificial Analysis launches new benchmark leaderboard for Wisedocs MLCR

What is measured here exactly

Artificial Analysis announced MLCR-AA, its own implementation of the MLCR (Medical Long Context Reasoning) benchmark from Wisedocs. The test ran only the private Expert and Compound layers, using text alone and no noisy documents, i.e., a model-only evaluation on a long medical record. The score combines three dimensions: completeness, accuracy and concision. According to the researchers, accuracy is the easier dimension, while completeness is where models fall short.

Top of the table

Claude Fable 5 from Anthropic achieved the highest score, 64.4%. The newer Claude models follow behind, but the price gap is large: cost per task ranges from $0.30 to $1 (≈1.1 to 3.7 shekel), depending on the model. That represents a gap of up to sixfold compared with K3 (Kimi K3), which sits just behind Claude on the Pareto front of score versus cost.

GPT-5.6, high accuracy, low completeness

OpenAI’s GPT-5.6 family does not reach the top scores because of low completeness, but it shows strong accuracy at relatively low cost. The two models in the family, Terra and Luna, sit on the Pareto front, meaning no other model offers higher accuracy for less money without sacrificing completeness. This is notable for users who need precise answers and are willing to fill in missing information manually.

What’s missing from this benchmark

Artificial Analysis explicitly states that it does not pad case files with filler documents, so the test measures reasoning on a “clean” long-context scenario rather than handling real-world noise. Performance metrics for open-source models or local runs were also not published. Bottom line: if you are building a product that must extract details from long medical records, Claude Fable 5 currently delivers the most complete result, but its cost can add up quickly in production.