GPT-6 automates 20.8% of remote-work projects, but that still sits near the floor

A jump from 2.5% to 20.8% automation of real projects sounds impressive, yet the new Remote Labor Index shows advanced models still fail the vast majority of economically valuable tasks.
What the index measures
The RLI, developed by researchers led by Mantas Mazeika, draws on actual projects pulled from the remote-work economy — game development, product design, architecture, data analysis, and video animation. Every project was priced and timed by the professionals who originally completed it. The dataset covers more than 6,000 work hours valued at over 140 thousand dollars; individual projects run up to $10,000 and 100-plus hours.
Results: measurable progress, low absolute performance
Last October the automation rate stood at just 2.5%. The new GPT-6 Astra version pushes that to 20.8%, but the researchers stress the figure still leaves models "close to the floor" — the overwhelming majority of projects were not completed at a level that would pass as commissioned work. They note the improvement is steady and measurable, giving the field a common baseline for tracking AI automation's trajectory.
Methodology and limits
The index evaluates end-to-end agent performance, not just knowledge or reasoning on academic benchmarks. Projects were randomly sampled and span a wide difficulty range. The paper appears as a preprint on arXiv (2510.26787) and has not yet been peer-reviewed; the numbers reflect the current version of the test framework.
What it means in practice
The gap between saturation on research benchmarks and failure on real economic tasks remains wide. The RLI offers a more grounded yardstick — not vendor promises, but projects someone actually paid for. For now, the trajectory points to gradual improvement, not a breakthrough that reshapes the labor market tomorrow.