HyperBrowseComp leaves 58% of multilingual search questions unsolved

Alham Fikri Aji's 423-question benchmark spans 13 languages and tests whether agents can connect evidence across web pages, video, maps and documents.

By · Published

Primary source: X

Why it matters

The benchmark exposes two operational constraints behind search-agent claims: strong results can depend on a provider's own search stack, and alternative retrieval setups can spend hundreds of millions of tokens while answering fewer questions. It also shows why multilingual scores need difficulty controls before they can support clean comparisons.

A magnifying glass crosses a mosaic of varied scripts and pictorial clues, evoking HyperBrowseComp’s multilingual, multimodal search-agent benchmark.

A new benchmark from Alham Fikri Aji (@AlhamFikri) found that five evaluated AI systems failed to answer 244 of its 423 questions using their built-in web search. The result, reported in the team's paper, puts a number on a stubborn weakness in search agents: finding and checking evidence when clues cross languages and media formats.…

Reader comments

Conversation for this story loads after sign-in.