Home
Academic Team Ends Tech Giants' Monopoly with SFT; OpenSeeker-v2 Tops Search Agent Rankings
In the evolving landscape of large language models (LLMs), deep search capabilities have become the decisive advantage for advanced intelligent agents. Yet this arena has long been controlled by well-funded industry giants. Conventional development approaches typically depend on resource-heavy pipelines that encompass pre-training, continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL).
A team of academic researchers has recently unveiled their latest work, OpenSeeker-v2, fundamentally challenging this established view. According to their report, training on high-quality, high-difficulty task trajectories enables even a straightforward supervised fine-tuning (SFT) approach to produce a top-tier search agent.

The team outlined three key optimization strategies for data synthesis: first, scaling up the knowledge graph to create a broader exploration space; second, substantially expanding the toolkit repertoire to push functional limits; and finally, applying strict low-step filtering to guarantee the quality and efficiency of training data.
Experimental results reveal that OpenSeeker-v2 (with 30B parameters and a ReAct architecture), trained on merely 10,600 data points, exhibited commanding performance across four core benchmarks: it scored 46.0% on BrowseComp, 58.1% on BrowseComp-ZH, 34.6% on "Humanity's Last Exam", and 78.0% on xbench. These figures not only set new records but also comprehensively outperformed industry models that relied on heavy pipelines combining CPT, SFT, and RL—such as Tongyi DeepResearch.

Strikingly, this marks the first state-of-the-art (SOTA) search agent developed by an entirely academic team using only SFT, at the same model scale and architecture. The team has now open-sourced the model weights for OpenSeeker-v2. This breakthrough significantly lowers the R&D barrier for advanced search agents and offers a more accessible, lightweight development pathway for both the academic and open-source communities.
Paper: https://arxiv.org/pdf/2605.04036
Related article
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
California AV Compliance: A New Era of Tickets, Geofences, and 1M Miles
Guident operates an AuveTech shuttle in South Florida, managing a four-mile route in West Palm Beach and a one-mile route in Boca Raton using its remote monitoring technology. | Credit: GuidentCalifornia is redefining the regulatory landscape for dri
Related Special Topic Recommendations
Comments (0)
0/500
In the evolving landscape of large language models (LLMs), deep search capabilities have become the decisive advantage for advanced intelligent agents. Yet this arena has long been controlled by well-funded industry giants. Conventional development approaches typically depend on resource-heavy pipelines that encompass pre-training, continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL).
A team of academic researchers has recently unveiled their latest work, OpenSeeker-v2, fundamentally challenging this established view. According to their report, training on high-quality, high-difficulty task trajectories enables even a straightforward supervised fine-tuning (SFT) approach to produce a top-tier search agent.

The team outlined three key optimization strategies for data synthesis: first, scaling up the knowledge graph to create a broader exploration space; second, substantially expanding the toolkit repertoire to push functional limits; and finally, applying strict low-step filtering to guarantee the quality and efficiency of training data.
Experimental results reveal that OpenSeeker-v2 (with 30B parameters and a ReAct architecture), trained on merely 10,600 data points, exhibited commanding performance across four core benchmarks: it scored 46.0% on BrowseComp, 58.1% on BrowseComp-ZH, 34.6% on "Humanity's Last Exam", and 78.0% on xbench. These figures not only set new records but also comprehensively outperformed industry models that relied on heavy pipelines combining CPT, SFT, and RL—such as Tongyi DeepResearch.

Strikingly, this marks the first state-of-the-art (SOTA) search agent developed by an entirely academic team using only SFT, at the same model scale and architecture. The team has now open-sourced the model weights for OpenSeeker-v2. This breakthrough significantly lowers the R&D barrier for advanced search agents and offers a more accessible, lightweight development pathway for both the academic and open-source communities.
Paper: https://arxiv.org/pdf/2605.04036
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage











