Microsoft unveils powerful new AI inference chip

Microsoft has unveiled its newest chip, the Maia 200, calling it a silicon workhorse built to scale AI inference.
Succeeding the Maia 100 launched in 2023, the Maia 200 is engineered to run advanced AI models with greater speed and efficiency. It packs over 100 billion transistors, achieving more than 10 petaflops in 4-bit precision and roughly 5 petaflops in 8-bit performance — a significant leap from its predecessor.
Inference is the computational process of running a model, as opposed to training it. As AI companies mature, inference costs now account for a growing share of total operating expenses, fueling renewed efforts to streamline the process.
Microsoft aims for the Maia 200 to support this optimization, helping AI businesses run with fewer disruptions and lower power usage. “In practical terms, a single Maia 200 node can effortlessly handle today’s largest models, with ample headroom for even bigger models in the future,” the company said.
Microsoft’s new chip also fits into a broader trend of tech giants developing their own silicon to cut reliance on Nvidia, whose advanced GPUs have become essential for AI success. Google, for instance, offers its TPU as cloud-accessible compute power rather than selling chips directly. Amazon has its Trainium AI accelerator, with the latest version, Trainium3, launched in December. These chips offload some computing tasks from Nvidia GPUs, reducing overall hardware expenses.
With Maia, Microsoft positions itself to compete with these alternatives. In its Monday press release, the company noted that Maia delivers three times the FP4 performance of third-generation Amazon Trainium chips and surpasses Google’s seventh-generation TPU in FP8 performance.
Microsoft reports that Maia is already powering the company’s AI models from its Superintelligence team and supporting its Copilot chatbot. As of Monday, the company has invited developers, academics, and frontier AI labs to use the Maia 200 software development kit in their workloads.
Techcrunch event
Disrupt 2026 Tickets: Exclusive One-Time Offer
Tickets are now available. Save up to $680 while these rates last, and be among the first 500 registrants to get 50% off your +1 pass. TechCrunch Disrupt brings together top leaders from Google Cloud, Netflix, Microsoft, Box, a16z, Hugging Face, and more across 250+ sessions designed to accelerate growth and sharpen your competitive edge. Connect with hundreds of innovative startups and join curated networking that drives deals, insights, and inspiration.
Disrupt 2026 Tickets: Exclusive One-Time Offer
Tickets are now available. Save up to $680 while these rates last, and be among the first 500 registrants to get 50% off your +1 pass. TechCrunch Disrupt brings together top leaders from Google Cloud, Netflix, Microsoft, Box, a16z, Hugging Face, and more across 250+ sessions designed to accelerate growth and sharpen your competitive edge. Connect with hundreds of innovative startups and join curated networking that drives deals, insights, and inspiration.
San Francisco | October 13-15, 2026 REGISTER NOW
Related article
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
California AV Compliance: A New Era of Tickets, Geofences, and 1M Miles
Guident operates an AuveTech shuttle in South Florida, managing a four-mile route in West Palm Beach and a one-mile route in Boca Raton using its remote monitoring technology. | Credit: GuidentCalifornia is redefining the regulatory landscape for dri
Related Special Topic Recommendations
Comments (1)
0/500
Just saw the news about Microsoft's new Maia 200 chip. Finally, someone is tackling the inference bottleneck with actual silicon instead of just throwing money at cloud services. 🚀 It’s wild how fast these models are evolving, but I’m still skeptical about whether this will actually lower costs for regular users or just make big tech richer. Anyone else excited for cheaper AI access, or is it all hype? 🤔

Microsoft has unveiled its newest chip, the Maia 200, calling it a silicon workhorse built to scale AI inference.
Succeeding the Maia 100 launched in 2023, the Maia 200 is engineered to run advanced AI models with greater speed and efficiency. It packs over 100 billion transistors, achieving more than 10 petaflops in 4-bit precision and roughly 5 petaflops in 8-bit performance — a significant leap from its predecessor.
Inference is the computational process of running a model, as opposed to training it. As AI companies mature, inference costs now account for a growing share of total operating expenses, fueling renewed efforts to streamline the process.
Microsoft aims for the Maia 200 to support this optimization, helping AI businesses run with fewer disruptions and lower power usage. “In practical terms, a single Maia 200 node can effortlessly handle today’s largest models, with ample headroom for even bigger models in the future,” the company said.
Microsoft’s new chip also fits into a broader trend of tech giants developing their own silicon to cut reliance on Nvidia, whose advanced GPUs have become essential for AI success. Google, for instance, offers its TPU as cloud-accessible compute power rather than selling chips directly. Amazon has its Trainium AI accelerator, with the latest version, Trainium3, launched in December. These chips offload some computing tasks from Nvidia GPUs, reducing overall hardware expenses.
With Maia, Microsoft positions itself to compete with these alternatives. In its Monday press release, the company noted that Maia delivers three times the FP4 performance of third-generation Amazon Trainium chips and surpasses Google’s seventh-generation TPU in FP8 performance.
Microsoft reports that Maia is already powering the company’s AI models from its Superintelligence team and supporting its Copilot chatbot. As of Monday, the company has invited developers, academics, and frontier AI labs to use the Maia 200 software development kit in their workloads.
Techcrunch eventDisrupt 2026 Tickets: Exclusive One-Time Offer
Tickets are now available. Save up to $680 while these rates last, and be among the first 500 registrants to get 50% off your +1 pass. TechCrunch Disrupt brings together top leaders from Google Cloud, Netflix, Microsoft, Box, a16z, Hugging Face, and more across 250+ sessions designed to accelerate growth and sharpen your competitive edge. Connect with hundreds of innovative startups and join curated networking that drives deals, insights, and inspiration.
Disrupt 2026 Tickets: Exclusive One-Time Offer
Tickets are now available. Save up to $680 while these rates last, and be among the first 500 registrants to get 50% off your +1 pass. TechCrunch Disrupt brings together top leaders from Google Cloud, Netflix, Microsoft, Box, a16z, Hugging Face, and more across 250+ sessions designed to accelerate growth and sharpen your competitive edge. Connect with hundreds of innovative startups and join curated networking that drives deals, insights, and inspiration.
San Francisco | October 13-15, 2026 REGISTER NOW
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
Just saw the news about Microsoft's new Maia 200 chip. Finally, someone is tackling the inference bottleneck with actual silicon instead of just throwing money at cloud services. 🚀 It’s wild how fast these models are evolving, but I’m still skeptical about whether this will actually lower costs for regular users or just make big tech richer. Anyone else excited for cheaper AI access, or is it all hype? 🤔





Home






