Meitun LongCat Launches LoHoSearch as BrowseComp Scores Drop Below 30%
Search agent capabilities have been largely defined by BrowseComp over the last year. Yet, this benchmark is losing its relevance — it drove model performance from 30% to 90% in just ten months, and its value is rapidly diminishing. On July 17, Meituan’s LongCat launched a more rigorous evaluation, LoHoSearch, designed to reignite meaningful progress in search agent development.

LoHoSearch’s challenges are not manually crafted but automatically generated from a Wikipedia knowledge graph containing 7.62 million entities. Because they are machine-generated, these tasks reach levels of search space and structural complexity that human annotators could never consistently replicate. In short, while previous benchmarks were limited by human-designed ceilings, LoHoSearch delegates question generation to the knowledge graph, effectively removing that upper bound.

The results are stark. Among 11 leading models, the highest score reached only 34.74%, with the next three scoring between 15% and 16%. By contrast, these same models achieve approximately 90% on BrowseComp. This gap highlights the true limitations of current search agents. Notably, context strategies yield diminishing returns: the best context approach on LoHoSearch improved scores by just 6.8 percentage points, compared to 14 percentage points on BrowseComp. This suggests that relying on prompts and context engineering to bridge capability gaps is largely ineffective under this new benchmark.
LoHoSearch comprises 544 questions spanning 11 domains, structured in tree and graph formats, and has been open-sourced. As older benchmarks plateau, the real challenges for search agents are just beginning, and LoHoSearch may soon become an essential evaluation standard.
Related article
Senspeech X2.5 Twin Stars: First Million-Token Context on the Edge, Fully Trained with Domestic Computing Power
On September 1st, iFLYTEK’s wholly-owned subsidiary, Ciyuan Xinghuo, officially launched and open-sourced two edge-side general large models: Xinghuo X2.5-4B and Xinghuo X2.5-1.7B. These models are the first of their kind to natively support a contex
MiniMax Unveils M3, a Domestic AI Large Model That Surpasses GPT-5.5
China’s AI sector has witnessed a major technological leap with the official launch of Xiyu Technology’s latest large language model, MiniMax M3. This advanced system combines state-of-the-art coding proficiency with support for an ultra-long context
India mandates caller-ID apps to share spam data with telecom operators
India has expanded its anti-spam regulations to mandate that caller-ID and call-management applications share user spam reports with telecom operators, a move that has led Truecaller, a prominent spam-blocking app provider, to label the policy as ant
Related Special Topic Recommendations
Comments (0)
0/500
Search agent capabilities have been largely defined by BrowseComp over the last year. Yet, this benchmark is losing its relevance — it drove model performance from 30% to 90% in just ten months, and its value is rapidly diminishing. On July 17, Meituan’s LongCat launched a more rigorous evaluation, LoHoSearch, designed to reignite meaningful progress in search agent development.

LoHoSearch’s challenges are not manually crafted but automatically generated from a Wikipedia knowledge graph containing 7.62 million entities. Because they are machine-generated, these tasks reach levels of search space and structural complexity that human annotators could never consistently replicate. In short, while previous benchmarks were limited by human-designed ceilings, LoHoSearch delegates question generation to the knowledge graph, effectively removing that upper bound.

The results are stark. Among 11 leading models, the highest score reached only 34.74%, with the next three scoring between 15% and 16%. By contrast, these same models achieve approximately 90% on BrowseComp. This gap highlights the true limitations of current search agents. Notably, context strategies yield diminishing returns: the best context approach on LoHoSearch improved scores by just 6.8 percentage points, compared to 14 percentage points on BrowseComp. This suggests that relying on prompts and context engineering to bridge capability gaps is largely ineffective under this new benchmark.
LoHoSearch comprises 544 questions spanning 11 domains, structured in tree and graph formats, and has been open-sourced. As older benchmarks plateau, the real challenges for search agents are just beginning, and LoHoSearch may soon become an essential evaluation standard.
Senspeech X2.5 Twin Stars: First Million-Token Context on the Edge, Fully Trained with Domestic Computing Power
On September 1st, iFLYTEK’s wholly-owned subsidiary, Ciyuan Xinghuo, officially launched and open-sourced two edge-side general large models: Xinghuo X2.5-4B and Xinghuo X2.5-1.7B. These models are the first of their kind to natively support a contex
MiniMax Unveils M3, a Domestic AI Large Model That Surpasses GPT-5.5
China’s AI sector has witnessed a major technological leap with the official launch of Xiyu Technology’s latest large language model, MiniMax M3. This advanced system combines state-of-the-art coding proficiency with support for an ultra-long context
India mandates caller-ID apps to share spam data with telecom operators
India has expanded its anti-spam regulations to mandate that caller-ID and call-management applications share user spam reports with telecom operators, a move that has led Truecaller, a prominent spam-blocking app provider, to label the policy as ant





Home






