Skip to content

OpenRouter AI Search Test: Iterating Search Queries More Important Than Switching Search Engines

Aug 15, 13:52

According to Scout Beating monitoring, OpenRouter has launched Web Search Benchmarks, specifically testing how AI-assisted web search should be configured. The initial batch covers Exa, Parallel, Perplexity, and on-board model search, paired with models such as GPT-5.6 Sol, Claude Opus 5, and DeepSeek V4 Flash, comparing accuracy, cost, and speed in 4 search benchmark groups.

The most significant result is: the number of search rounds is more important than switching search engines. In challenging tasks like BrowseComp, increasing the search limit from 1 round to 25 rounds can approximately double the score, while the cost per question increases by about 2.5 to 7 times. For example, pairing Claude Opus 5 with Perplexity increases from 35.8% in 1 round to 89.0% in 25 rounds.

The model itself is also more critical than the search engine. Keeping other conditions constant, switching models on average can lead to a 15-point difference, while switching search engines like Exa, Parallel, Perplexity has an average impact of around 10 points. The model manufacturer's own search may not always be the best; it depends on the specific task.

However, searching more does not necessarily mean more cost-effective. In BrowseComp, on average, the model searches 10.3 times when answered correctly and 19.7 times when answered incorrectly. For tasks where the answers are already challenging to find, increasing the search frequency may just make the Agent more expensively fail.

Source