Liquid AI open-sources hybrid expert model LFM2.5 for edge AI
AI startup Liquid AI today officially released and open-sourced its latest edge-side large model, LFM2.5-8B-A1B . Designed for tool calling and complex instruction following on consumer hardware, the model boosts reasoning and inference performance on edge devices while keeping computational costs extremely low.
Architecturally, it uses a sparse mixture-of-experts (MoE) design with 8.3 billion total parameters. Thanks to that sparsity, it activates only 1.5 billion parameters per token, enabling smooth local operation on smartphones, laptops, and similar devices.

Extended Context and Stronger Reasoning
Compared to its predecessor, LFM2.5 extends the context window from 32K to 128K tokens and increases pre-training data from 12T to 38T tokens. As a pure inference model, it produces an explicit reasoning chain before delivering the final answer. Its highly compressed vocabulary efficiently handles nine languages, including Chinese and Arabic.
To reduce logical dead loops and hallucinations during long reasoning, the development team applied two-stage reinforcement learning (RL) during training. Preference optimization effectively cuts down on "dead loops" in long-chain reasoning, while a specialized anti-hallucination reward mechanism lets the model actively refuse to answer questions outside its knowledge base.
Powerful Edge Performance and Broad Ecosystem Support
In performance, LFM2.5 has seen explosive gains. Its scores on logical reasoning and anti-hallucination benchmarks far surpass its predecessor, and it even competes with larger models in instruction following. For tool calling, the model defaults to outputting efficient Python function calls and supports seamless switching to JSON format via system prompts.
The model received full support from major inference frameworks on launch day, including llama.cpp, MLX, vLLM, and SGLang. In hardware tests, it reached decoding speeds of up to 253 bytes per second on the M5 Max chip and about 30 bytes per second on mobile devices, perfectly balancing privacy and efficiency for edge deployments.
Related article
Anthropic Expands Claude AI Coding Tools to Japan in Push for Overseas Growth
Anthropic, a prominent U.S. artificial intelligence firm, is intensifying its global outreach. On Wednesday, the company hosted a major developer gathering, "Code with Claude," in Tokyo, drawing close to 500 software engineers. This initiative seeks
India’s Software Giant Cuts Hiring, Promises No Layoffs as AI Agents Scale Up
As artificial intelligence reshapes traditional labor-intensive business models, Tata Consultancy Services (TCS), a premier Indian software outsourcing firm, has unveiled its strategic response. During Tuesday’s annual general meeting, the chairman o
China Locks AI Models During Gaokam to Block Instant Homework Help
With the 2026 Gaokao fast approaching, rumors regarding the suspension of AI tools during the exam period have ignited intense online debate. In response to public concern, major AI platforms and educational apps have clarified their stance: rather t
Related Special Topic Recommendations
Comments (1)
0/500
AI startup Liquid AI today officially released and open-sourced its latest edge-side large model,
Architecturally, it uses a sparse mixture-of-experts (MoE) design with 8.3 billion total parameters. Thanks to that sparsity, it activates only 1.5 billion parameters per token, enabling smooth local operation on smartphones, laptops, and similar devices.

Extended Context and Stronger Reasoning
Compared to its predecessor, LFM2.5 extends the context window from 32K to 128K tokens and increases pre-training data from 12T to 38T tokens. As a pure inference model, it produces an explicit reasoning chain before delivering the final answer. Its highly compressed vocabulary efficiently handles nine languages, including Chinese and Arabic.
To reduce logical dead loops and hallucinations during long reasoning, the development team applied two-stage reinforcement learning (RL) during training. Preference optimization effectively cuts down on "dead loops" in long-chain reasoning, while a specialized anti-hallucination reward mechanism lets the model actively refuse to answer questions outside its knowledge base.
Powerful Edge Performance and Broad Ecosystem Support
In performance, LFM2.5 has seen explosive gains. Its scores on logical reasoning and anti-hallucination benchmarks far surpass its predecessor, and it even competes with larger models in instruction following. For tool calling, the model defaults to outputting efficient Python function calls and supports seamless switching to JSON format via system prompts.
The model received full support from major inference frameworks on launch day, including llama.cpp, MLX, vLLM, and SGLang. In hardware tests, it reached decoding speeds of up to 253 bytes per second on the M5 Max chip and about 30 bytes per second on mobile devices, perfectly balancing privacy and efficiency for edge deployments.
Anthropic Expands Claude AI Coding Tools to Japan in Push for Overseas Growth
Anthropic, a prominent U.S. artificial intelligence firm, is intensifying its global outreach. On Wednesday, the company hosted a major developer gathering, "Code with Claude," in Tokyo, drawing close to 500 software engineers. This initiative seeks
India’s Software Giant Cuts Hiring, Promises No Layoffs as AI Agents Scale Up
As artificial intelligence reshapes traditional labor-intensive business models, Tata Consultancy Services (TCS), a premier Indian software outsourcing firm, has unveiled its strategic response. During Tuesday’s annual general meeting, the chairman o
China Locks AI Models During Gaokam to Block Instant Homework Help
With the 2026 Gaokao fast approaching, rumors regarding the suspension of AI tools during the exam period have ignited intense online debate. In response to public concern, major AI platforms and educational apps have clarified their stance: rather t





Home






