Will it run?
Research

Nvidia study finds open language models lose track in long repetitive tasks

By Rae Whitlock Clawpit staff

A new Nvidia paper shows that open-weight language models become markedly less reliable as tasks grow longer, even when the full context fits inside the model's window. Researchers tested seven models on deliberately simple, repetitive chores — adding numbers, sorting lists, updating fields — and measured an average 62.8% drop in accuracy when moving from 4,000-token tasks to 128,000-token tasks. The best-performing model completed every item without error in only 17.1% of the longest runs.

The problem is not model size, the authors say, but tracking. The models "understood" the task and knew what to do, yet they lost their place in the list, especially when items lacked unique identifiers. The failure resembles an office worker handed a massive invoice file who skips a line or updates the wrong record, despite being able to read the entire file at once.

The experiment used seven open-weight models on synthetic tasks designed to isolate long-context tracking from semantic complexity. The gap between 4K and 128K tokens was consistent across every model. The standout figure — 17.1% full success on the leading model — means that in the vast majority of long runs, at least one item fell through the cracks.

The practical takeaway for agent developers is direct: if your agent processes long lists, give every item a unique identifier, break the work into small batches, and verify every line of output. This is not a matter of the model being insufficiently smart; it is a structural failure in state tracking over long context, and it is not solved by expanding the context window alone. The paper is available on arXiv as 2609.38712.