Home
Microsoft Bing Team Open Sources 27B Embedding Model Harrier, Leading Multilingual Benchmarks
On April 7, Microsoft's Bing team announced the open-source release of a new word embedding model series called "Harrier," designed to reshape the underlying logic of global search, retrieval, and AI agents. The Harrier series includes three versions, with the flagship 27B model surpassing major proprietary models from OpenAI, Amazon, and Google Gemini in the multilingual MTEB v2 benchmark, securing the top spot.

The technical foundation of this model reflects strong industrial standards: Harrier supports over 100 languages with a context window of up to 32,000 tokens. For training, Microsoft used over 2 billion real-world examples and incorporated synthetic data from GPT-5 for reinforcement learning. This high-quality data mix gives Harrier a distinct advantage in understanding complex contexts and handling long texts. Alongside the full 27B-parameter version, Microsoft also released smaller 0.6B and 2.7B variants to accommodate different computing environments, all open-sourced on Hugging Face under the MIT license.
Embedding models are essential for organizing and retrieving information in AI systems, and their performance directly impacts the accuracy of RAG (Retrieval-Augmented Generation) systems. Microsoft intends to integrate this technology into the Bing search engine and new AI agent grounding services. As AI moves toward more autonomous, multi-step tasks, the open-sourcing of Harrier not only gives developers a high-performance alternative to proprietary models but also marks a significant milestone for the open-source ecosystem in semantic representation, further accelerating AI agent deployment in multilingual global environments.
Related article
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
California AV Compliance: A New Era of Tickets, Geofences, and 1M Miles
Guident operates an AuveTech shuttle in South Florida, managing a four-mile route in West Palm Beach and a one-mile route in Boca Raton using its remote monitoring technology. | Credit: GuidentCalifornia is redefining the regulatory landscape for dri
Related Special Topic Recommendations
Comments (1)
0/500
Finally, an open-source embedding model that actually challenges the big players! Harrier's 27B parameters look promising for multilingual tasks, especially if it improves retrieval accuracy without the usual proprietary lock-in. Microsoft is pushing boundaries here, but let's see how it performs in real-world edge cases. Excited to test it out on some niche datasets! 🚀
On April 7, Microsoft's Bing team announced the open-source release of a new word embedding model series called "Harrier," designed to reshape the underlying logic of global search, retrieval, and AI agents. The Harrier series includes three versions, with the flagship 27B model surpassing major proprietary models from OpenAI, Amazon, and Google Gemini in the multilingual MTEB v2 benchmark, securing the top spot.

The technical foundation of this model reflects strong industrial standards: Harrier supports over 100 languages with a context window of up to 32,000 tokens. For training, Microsoft used over 2 billion real-world examples and incorporated synthetic data from GPT-5 for reinforcement learning. This high-quality data mix gives Harrier a distinct advantage in understanding complex contexts and handling long texts. Alongside the full 27B-parameter version, Microsoft also released smaller 0.6B and 2.7B variants to accommodate different computing environments, all open-sourced on Hugging Face under the MIT license.
Embedding models are essential for organizing and retrieving information in AI systems, and their performance directly impacts the accuracy of RAG (Retrieval-Augmented Generation) systems. Microsoft intends to integrate this technology into the Bing search engine and new AI agent grounding services. As AI moves toward more autonomous, multi-step tasks, the open-sourcing of Harrier not only gives developers a high-performance alternative to proprietary models but also marks a significant milestone for the open-source ecosystem in semantic representation, further accelerating AI agent deployment in multilingual global environments.
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
Finally, an open-source embedding model that actually challenges the big players! Harrier's 27B parameters look promising for multilingual tasks, especially if it improves retrieval accuracy without the usual proprietary lock-in. Microsoft is pushing boundaries here, but let's see how it performs in real-world edge cases. Excited to test it out on some niche datasets! 🚀











