Will it run?
Models

New training method targets LLM slop with expert-aligned rubrics

By Rae Whitlock Clawpit staff
New training method targets LLM slop with expert-aligned rubrics

Researchers have introduced a training approach that claims to solve the "slop" problem — the mediocre text language models produce when no single correct answer exists.

The problem: non-expert judges. Large language models reach superhuman performance on verifiable tasks such as code, mathematics and logic, but remain weak at open-ended writing. The issue originates in classic RLHF: reward models are trained on rankings from non-expert annotators, inheriting their ceiling. The quality that matters most — expert-level, even superhuman writing — is precisely what these judges are least equipped to recognize.

The experiment: academic section writing. The team built a concrete benchmark: a model receives a real academic paper minus one section — abstract, introduction, literature review or conclusion — and must write the missing piece to match the rest of the paper in length and style. The authors' original text serves as a natural expert reference, and the task is reproducible at scale from published literature. Two models served as both writers and judges: Opus-4.8 and GPT-5.6-sol, cross-evaluated against each other.

The surprising finding: the model beats the expert. A direct pairwise judge — presented with two candidates and asked which is better — preferred the model over the human expert in a majority of cases, between 63.5% and 84.6% depending on configuration. When the team switched to standard rubrics, GPT-5.6 built a meta-prompt that generated dedicated rubrics for each section, and the result intensified further: the model was selected as higher-scoring in 100% of cases. The judges, in short, did not detect the gap.

The solution: meta-optimization of rubrics. RL-XAR (Reinforcement Learning from eXpert-Aligned Rubrics) reverses the direction. The method generates rubrics, checks whether they rank expert writing above model writing, then updates the rubrics themselves — meta-optimization — until the gap is no longer detectable. Only then are the aligned rubrics used for reinforcement training. The process repeats until meta-optimization finds no further significant gap.

Preliminary results and what is missing. According to the publication, multiple metrics show large improvement over standard training on three tasks: academic section writing, continuations of Pulitzer-winning novels, and high-quality Wikipedia pages. The paper does not publish full benchmark numbers, does not specify sample sizes for each task, and does not state whether results have undergone peer review — it is a preprint. The promise is large; the proof remains partial.