Will it run?
Models

OpenRouter launches network search benchmark: search budget outweighs engine

By Rae Whitlock Clawpit staff
OpenRouter launches network search benchmark: search budget outweighs engine

OpenRouter released a new suite of benchmarks that rank search tools across different models and configurations, helping developers decide how to ground their agents. The study covered four diverse benchmarks, four models, four search depths and four engines—Exa, Parallel, Perplexity and each model’s native engine. All combinations were evaluated on quality, cost and speed.

The most striking finding is that increasing the search budget had the largest impact on results. Moving from a budget of one tranche to 25 tranches roughly doubled the scores on the BrowseComp benchmark. In production, paying for more search rounds yields significantly more accurate answers.

Switching the model produced an average gap of about 15 points between frontier models and cost-efficient models. Changing the search engine while keeping the same model shifted scores by only about 10 points. In other words, investing in a stronger model returns more benefit than swapping the search provider.

In the BrowseComp benchmark with a 25-tranche budget, a cost gap emerged: successful runs averaged 10.3 searches, while failed runs averaged 19.7 searches. If a high failure rate is expected, reducing the number of tranches can cut costs without markedly harming successful outcomes.

The test also disproved the assumption that a model’s native engine is always the best default. For example, the native search of GPT 5.6 Sol beat a third-party engine in only 50 % of the benchmarks. The optimal choice depends on the specific workload; there is no one-size-fits-all solution.