Tech companies warm to cost-effective AI models

The AI boom has relied on a core assumption: larger models are more capable, and the most capable models win. Now the industry is about to discover what happens when that assumption begins to crack.
Mounting costs have already pressured users to reconsider smaller, cheaper models. This cost-conscious approach to model selection is new, and it's unclear how it will reshape the industry — but the impact is likely to be significant.
One prediction, articulated most clearly by Coinbase co-founder Brian Armstrong, is that the vast majority of tasks will shift to cheaper models.
“Demand for intelligence is near infinite, but 80% of workloads will run on models that are 99% cheaper within 12 to 18 months,” Armstrong wrote on X. “20% of workloads will still run on the latest generation models where maximizing IQ is important.”
It's hard to overstate how significant a shift this would be for the AI industry if Armstrong's prediction comes true.
Before now, most AI companies have competed on quality, which has meant defaulting to the most advanced model available. If those same tasks can be handled by cheaper models without any loss in quality, it would signal a massive shift in the economics of AI. And critically, much of the savings would come directly from the big labs' revenues — dealing a financial blow to OpenAI and Anthropic just as they head toward their IPOs.
This is a potentially seismic change for the industry, hinging on one fundamental question: Are companies ready to switch to smaller models?
Initial tests suggest that, with the right setup, cheaper models can step in without any sacrifice in quality. In a recent test by the legal AI tool Harvey, the company was able to reduce inference costs by three times without lowering quality. The test, conducted in partnership with the inference platform Fireworks AI, combined Claude Opus and Fireworks' GLM 5.1, and switched to Opus for the most demanding tasks. The result was a significantly lighter load on servers and overall lower costs.
“Quality comes first, and in legal it always will,” Harvey co-founder Gabe Pereyra told TechCrunch, referring to his startup's AI legal services. “However, the definition of quality is evolving from simply using the most powerful model for everything, to using the best model that gets the right answer most efficiently.”
This trend is often framed as major labs versus Chinese models or open-weight ones, but that misses the bigger picture. The real divide isn't between proprietary and open models — it's between large models and small ones. You can save money by switching from GPT-5.5 to DeepSeek's V4Flash, but switching to GPT-5.4-mini works just as well.
There's an active price war between in-house inference from the big labs and independently served open-weight models. For the larger question of small versus large, it doesn't really matter which kind of small model wins.
All of this might seem obvious — of course you shouldn't use more compute than necessary — but it runs counter to the scaling-first approach that has dominated the industry until now. Inspired by the bitter lesson, labs have leaned heavily into training the most compute-intensive models possible, pushing the frontier of what AI models can do. With prices heavily subsidized by investors, clients had no reason to choose anything but the most advanced option.
With token prices rising and subsidies slowing, users are facing cost pressure for the first time. We don't yet know whether this new cost pressure will actually drive enterprise users toward smaller models. They could just as easily economize by making fewer calls, using shorter context windows, or simply abandoning their least promising deployments.
But if it turns out that most deployments can run just as well on a smaller model, it could put a serious damper on the growing demand for inference — and raise new questions about how to justify the cost of training a frontier model.
Related article
Anthropic launches Opus 4.8 featuring new dynamic workflow tool
Anthropic unveiled Opus 4.8 on Thursday, marking the latest iteration of its premier public model. Priced identically to its predecessor, this update is now accessible across all platforms.Releasing just 41 days after Opus 4.7, Anthropic has accelera
OpenAI launches safer ChatGPT for teens years after they started using it
Following a series of lawsuits regarding the absence of safety protocols in AI chatbots—which contributed to teen suicides and other mental health crises—OpenAI unveiled ChatGPT for Teens on Monday. This new offering incorporates enhanced safety feat
Frontier AI Labs Refuse to Disclose Containment Strategies for Rogue Models
Recent research indicates that very few leading AI laboratories have published or demonstrated containment response plans. A containment plan defines the procedures for when an AI system attempts to subvert human control, specifying which access righ
Related Special Topic Recommendations
Comments (0)
0/500

The AI boom has relied on a core assumption: larger models are more capable, and the most capable models win. Now the industry is about to discover what happens when that assumption begins to crack.
Mounting costs have already pressured users to reconsider smaller, cheaper models. This cost-conscious approach to model selection is new, and it's unclear how it will reshape the industry — but the impact is likely to be significant.
One prediction, articulated most clearly by Coinbase co-founder Brian Armstrong, is that the vast majority of tasks will shift to cheaper models.
“Demand for intelligence is near infinite, but 80% of workloads will run on models that are 99% cheaper within 12 to 18 months,” Armstrong wrote on X. “20% of workloads will still run on the latest generation models where maximizing IQ is important.”
It's hard to overstate how significant a shift this would be for the AI industry if Armstrong's prediction comes true.
Before now, most AI companies have competed on quality, which has meant defaulting to the most advanced model available. If those same tasks can be handled by cheaper models without any loss in quality, it would signal a massive shift in the economics of AI. And critically, much of the savings would come directly from the big labs' revenues — dealing a financial blow to OpenAI and Anthropic just as they head toward their IPOs.
This is a potentially seismic change for the industry, hinging on one fundamental question: Are companies ready to switch to smaller models?
Initial tests suggest that, with the right setup, cheaper models can step in without any sacrifice in quality. In a recent test by the legal AI tool Harvey, the company was able to reduce inference costs by three times without lowering quality. The test, conducted in partnership with the inference platform Fireworks AI, combined Claude Opus and Fireworks' GLM 5.1, and switched to Opus for the most demanding tasks. The result was a significantly lighter load on servers and overall lower costs.
“Quality comes first, and in legal it always will,” Harvey co-founder Gabe Pereyra told TechCrunch, referring to his startup's AI legal services. “However, the definition of quality is evolving from simply using the most powerful model for everything, to using the best model that gets the right answer most efficiently.”
This trend is often framed as major labs versus Chinese models or open-weight ones, but that misses the bigger picture. The real divide isn't between proprietary and open models — it's between large models and small ones. You can save money by switching from GPT-5.5 to DeepSeek's V4Flash, but switching to GPT-5.4-mini works just as well.
There's an active price war between in-house inference from the big labs and independently served open-weight models. For the larger question of small versus large, it doesn't really matter which kind of small model wins.
All of this might seem obvious — of course you shouldn't use more compute than necessary — but it runs counter to the scaling-first approach that has dominated the industry until now. Inspired by the bitter lesson, labs have leaned heavily into training the most compute-intensive models possible, pushing the frontier of what AI models can do. With prices heavily subsidized by investors, clients had no reason to choose anything but the most advanced option.
With token prices rising and subsidies slowing, users are facing cost pressure for the first time. We don't yet know whether this new cost pressure will actually drive enterprise users toward smaller models. They could just as easily economize by making fewer calls, using shorter context windows, or simply abandoning their least promising deployments.
But if it turns out that most deployments can run just as well on a smaller model, it could put a serious damper on the growing demand for inference — and raise new questions about how to justify the cost of training a frontier model.
Anthropic launches Opus 4.8 featuring new dynamic workflow tool
Anthropic unveiled Opus 4.8 on Thursday, marking the latest iteration of its premier public model. Priced identically to its predecessor, this update is now accessible across all platforms.Releasing just 41 days after Opus 4.7, Anthropic has accelera
OpenAI launches safer ChatGPT for teens years after they started using it
Following a series of lawsuits regarding the absence of safety protocols in AI chatbots—which contributed to teen suicides and other mental health crises—OpenAI unveiled ChatGPT for Teens on Monday. This new offering incorporates enhanced safety feat
Frontier AI Labs Refuse to Disclose Containment Strategies for Rogue Models
Recent research indicates that very few leading AI laboratories have published or demonstrated containment response plans. A containment plan defines the procedures for when an AI system attempts to subvert human control, specifying which access righ





Home






