AI Scaling Breakthrough Questioned by Experts

There's been some buzz on social media about researchers discovering a new AI "scaling law," but experts are taking it with a grain of salt. AI scaling laws, which are more like informal guidelines, show how AI models get better as you throw more data and computing power at them. Up until about a year ago, the big trend was all about "pre-training" – basically, training bigger models on bigger datasets. That's still a thing, but now we've got two more scaling laws in the mix: post-training scaling, which is all about tweaking a model's behavior, and test-time scaling, which involves using more computing power during inference to boost a model's "reasoning" capabilities (think models like R1).
Recently, researchers from Google and UC Berkeley dropped a paper that some folks online are calling a fourth law: "inference-time search." This method has the model spit out a bunch of possible answers to a query at the same time and then pick the best one. The researchers claim it can juice up the performance of an older model, like Google's Gemini 1.5 Pro, to beat OpenAI's o1-preview "reasoning" model on science and math benchmarks.
Eric Zhao, a Google doctorate fellow and one of the paper's co-authors, shared on X that by just randomly sampling 200 responses and letting the model self-verify, Gemini 1.5 – which he jokingly called an "ancient early 2024 model" – could outdo o1-preview and even get close to o1. He pointed out that self-verification gets easier as you scale up, which is kind of counterintuitive but cool.
But not everyone's convinced. Matthew Guzdial, an AI researcher and assistant professor at the University of Alberta, told TechCrunch that this approach works best when you've got a solid way to judge the answers. Most questions aren't that straightforward, though. He said, "If we can't write code to define what we want, we can't use [inference-time] search. For something like general language interaction, we can't do this... It's generally not a great approach to actually solving most problems."
Zhao responded, saying their paper actually looks at cases where you don't have a clear way to judge the answers, and the model has to figure it out on its own. He argued that the gap between having a clear way to judge and not having one can shrink as you scale up.
Mike Cook, a research fellow at King's College London, backed up Guzdial's view, saying that inference-time search doesn't really make the model's reasoning better. It's more like a workaround for the model's tendency to make confident mistakes. He pointed out that if your model messes up 5% of the time, checking 200 attempts should make those mistakes easier to spot.
This news might be a bit of a downer for the AI industry, which is always on the hunt for ways to boost model "reasoning" without breaking the bank. As the paper's authors noted, reasoning models can rack up thousands of dollars in computing costs just to solve one math problem.
Looks like the search for new scaling techniques is far from over.
*Updated 3/20 5:12 a.m. Pacific: Added comments from study co-author Eric Zhao, who takes issue with an assessment by an independent researcher who critiqued the work.*
Related article
Bayer Leverages Iambic AI to Speed Drug Discovery
Juergen Eckhardt, M.D., serves as Head of Business Development and Licensing at Bayer Pharmaceuticals.Bayer is leveraging Iambic Therapeutics to deploy frontier AI models for drug discovery, alongside a major decarbonization project at its plant in S
Optimization-Driven AI Emerges as New Path to General-Purpose Models
Researchers from the University of Illinois Urbana-Champaign and the University of Virginia have created a new model architecture that could pave the way for more resilient AI systems with enhanced reasoning power.Named the energy-based transformer (
AI Boom Echoes Dot-Com Era Bubble Concerns
The influx of multi-billion dollar investments into AI has fueled a heated debate: is the industry headed for a dot-com style bubble?Investors are vigilant for any cooling of enthusiasm or signs that massive spending on chips and infrastructure isn't
Related Special Topic Recommendations
Comments (37)
0/500
I heard there's a new AI scaling law discovered, but experts are skeptical? It feels like another case of hype overshadowing reality—maybe we should wait for solid experiments before jumping to conclusions. 😅
Interessant, aber ich bin skeptisch. Diese 'Skalierungsgesetze' klingen oft nach einer selbsterfüllenden Prophezeiung der großen Tech-Firmen. Mehr Daten, mehr Rechenleistung – klar wird das Modell 'besser', aber zu welchem Preis? Die Umweltkosten sind enorm, und am Ende bekommen wir vielleicht nur bessere Halluzinationen. Die Experten haben recht, vorsichtig zu sein. 🤔
This AI scaling law thing sounds cool, but it's hard to get excited when experts are so skeptical. It's like they're saying, 'Sure, it's interesting, but let's not get carried away.' I guess we'll see if it's the real deal or just another hype train. 🤔
Essa história de lei de escalabilidade de IA parece legal, mas é difícil se empolgar quando os especialistas são tão céticos. Parece que eles estão dizendo, 'Sim, é interessante, mas não vamos nos empolgar muito'. Vamos ver se é verdade ou só mais um hype. 🤔

Bayer Leverages Iambic AI to Speed Drug Discovery
Juergen Eckhardt, M.D., serves as Head of Business Development and Licensing at Bayer Pharmaceuticals.Bayer is leveraging Iambic Therapeutics to deploy frontier AI models for drug discovery, alongside a major decarbonization project at its plant in S
Optimization-Driven AI Emerges as New Path to General-Purpose Models
Researchers from the University of Illinois Urbana-Champaign and the University of Virginia have created a new model architecture that could pave the way for more resilient AI systems with enhanced reasoning power.Named the energy-based transformer (
AI Boom Echoes Dot-Com Era Bubble Concerns
The influx of multi-billion dollar investments into AI has fueled a heated debate: is the industry headed for a dot-com style bubble?Investors are vigilant for any cooling of enthusiasm or signs that massive spending on chips and infrastructure isn't
I heard there's a new AI scaling law discovered, but experts are skeptical? It feels like another case of hype overshadowing reality—maybe we should wait for solid experiments before jumping to conclusions. 😅
Interessant, aber ich bin skeptisch. Diese 'Skalierungsgesetze' klingen oft nach einer selbsterfüllenden Prophezeiung der großen Tech-Firmen. Mehr Daten, mehr Rechenleistung – klar wird das Modell 'besser', aber zu welchem Preis? Die Umweltkosten sind enorm, und am Ende bekommen wir vielleicht nur bessere Halluzinationen. Die Experten haben recht, vorsichtig zu sein. 🤔
This AI scaling law thing sounds cool, but it's hard to get excited when experts are so skeptical. It's like they're saying, 'Sure, it's interesting, but let's not get carried away.' I guess we'll see if it's the real deal or just another hype train. 🤔
Essa história de lei de escalabilidade de IA parece legal, mas é difícil se empolgar quando os especialistas são tão céticos. Parece que eles estão dizendo, 'Sim, é interessante, mas não vamos nos empolgar muito'. Vamos ver se é verdade ou só mais um hype. 🤔





Home






