option
Home
News
Anthropic Claims AI Isn't Stalling, It's Outsmarting Benchmarks

Anthropic Claims AI Isn't Stalling, It's Outsmarting Benchmarks

April 17, 2025
181

Anthropic Claims AI Isn

Large language models (LLMs) and other generative AI technologies are making significant strides in self-correction, which is paving the way for new applications, including what's known as "agentic AI," according to Michael Gerstenhaber, Vice President of Anthropic, a leading AI model developer.

"It's getting very good at self-correction, self-reasoning," Gerstenhaber, who leads API technologies at Anthropic, shared during an interview in New York with Bloomberg Intelligence's Anurag Rana. Anthropic, creators of the Claude family of LLMs, are direct competitors to OpenAI's GPT models. "Every couple of months, we release a new model that expands the capabilities of LLMs," he added, emphasizing the dynamic nature of the industry where each model revision unlocks new potential uses.

New Capabilities in AI Models

The latest models from Anthropic have introduced capabilities such as task planning, allowing them to perform tasks on a computer much like a human would, like ordering pizza online. "Planning interstitial steps, something that wasn't feasible yesterday, is now within reach," Gerstenhaber noted about this step-by-step task execution.

The discussion, which also featured Vijay Karunamurthy, Chief Technologist at AI startup Scale AI, was part of a daylong conference hosted by Bloomberg Intelligence titled "Gen AI: Can it deliver on the productivity promise?"

Challenging AI Skepticism

Gerstenhaber's insights challenge the views of AI skeptics who argue that generative AI and the broader AI field are "hitting a wall," suggesting diminishing returns with each new model iteration. AI scholar Gary Marcus, for instance, has been vocal about his concerns since 2022, warning that simply increasing the size of AI models (more parameters) won't proportionally improve their performance.

However, Gerstenhaber asserts that Anthropic is pushing the boundaries beyond what current AI benchmarks can measure. "Even if it looks like progress is slowing in some areas, it's because we're unlocking entirely new functionalities, but we've saturated the benchmarks and the ability to perform older tasks," he explained. This makes it increasingly difficult to gauge the full extent of what current generative AI models can achieve.

Scaling and Learning

Both Gerstenhaber and Karunamurthy emphasized the importance of scaling generative AI models to enhance their self-correcting capabilities. "We're definitely seeing more and more scaling of the intelligence," Gerstenhaber remarked. Karunamurthy added, "One reason we believe we're not hitting a wall with planning and reasoning is that we're still learning how to structure these tasks so that the models can adapt to new and varied environments."

Gerstenhaber agreed, stating, "We're in the early stages, learning from application developers about their needs and where the models fall short, which we can then integrate back into the language model."

Real-Time Learning and Adaptation

Much of this progress, according to Gerstenhaber, is driven by the rapid pace of fundamental research at Anthropic, as well as real-time learning from industry feedback. "We're adapting to what the industry tells us they need, learning in real time," he said.

Customers often start with larger models and then scale down to simpler ones to suit specific purposes. "Initially, they assess whether a model is intelligent enough to perform a task well, then whether it's fast enough to meet their application needs, and finally, if it can be as cost-effective as possible," Gerstenhaber explained.

Related article
Apple Smart Glasses Could Debut at WWDC27, Highlighting Privacy Protection Apple Smart Glasses Could Debut at WWDC27, Highlighting Privacy Protection Bloomberg’s Mark Gurman reports that Apple’s smart glasses, codenamed N50, are slated for a WWDC27 debut in June 2027, with a retail launch expected in autumn 2027. Originally targeted for late this year and early 2027, the device’s release has been
Inside Details Exposed About Next-Gen Gemini: Strained Computing Power, Internal Teams Disagreed on Development Priorities and Resource Allocation Inside Details Exposed About Next-Gen Gemini: Strained Computing Power, Internal Teams Disagreed on Development Priorities and Resource Allocation Reports indicate that the launch of Google’s highly anticipated next-generation Gemini model has been pushed back. Internal disagreements over development priorities and resource allocation, combined with limited computing capacity and complex approv
OpenAI Dismisses Growth Slowdown Concerns, Says Multiple Business Units Accelerating OpenAI Dismisses Growth Slowdown Concerns, Says Multiple Business Units Accelerating In response to external scrutiny regarding decelerating sales growth and missed internal benchmarks, AI leader OpenAI issued a confident statement on Tuesday, April 28. The company clarified that its consumer products and enterprise services are adva
Related Special Topic Recommendations
Data Analysis AI SQL Copilots for Revenue Dashboards, Funnel Analysis, and Product Metrics
AI SQL Copilots for Revenue Dashboards, Funnel Analysis, and Product Metrics

2026 Latest Best AI SQL Copilots Ranked Top-Rated! XIX.AI curates a powerful game-changing collection for weekly updated real-world tests. These must-try tools help you generate accurate revenue dashboards, analyze sales funnels, and track product metrics swiftly, boosting productivity massively. Explore now to Discover your perfect tool for data-driven decision making! 238 characters

9 tools
xix.ai
Music composition Best AI Melody Writing Tools for Song Drafts
Best AI Melody Writing Tools for Song Drafts

2026 Latest Best Top-Rated AI Melody Writing Tools for Song Drafts! XIX.AI has curated a highly powerful game-changing collection that goes through rigorous real-world tests to deliver the best writing experience. You can find detailed free vs paid comparisons, accurate rankings, and must-try options designed to help you create stunning song drafts effortlessly and boost your creative productivity significantly. Explore now to discover your perfect tool!

8 tools
xix.ai
chatbot Best AI Conversation Trainer Tools for Interview Practice
Best AI Conversation Trainer Tools for Interview Practice

2026 Latest Best Top-rated AI Conversation Trainer Tools for Interview Practice are here on XIX.AI! This curated collection features powerful, game-changing tools that go through rigorous real-world tests to deliver accurate feedback. You’ll find a free vs paid comparison and detailed rankings to help you choose the must-try option that boosts your confidence and skills. Explore now to Discover your perfect tool for interview success!

12 tools
xix.ai
Design & Art Best AI Style Transfer Tools for Creative Experiments
Best AI Style Transfer Tools for Creative Experiments

2026 Latest Best Top-rated AI Style Transfer Tools for Creative Experiments! XIX.AI has curated a powerful, game-changing collection of must-try tools that deliver exceptional results through real-world tests and rigorous rankings. These top solutions help creatives boost productivity significantly by accelerating content creation and unlocking endless creative possibilities. Explore now to discover your perfect tool and start creating today!

9 tools
xix.ai
Comic Creation Best AI Dialogue Bubble Tools for Visual Storytelling
Best AI Dialogue Bubble Tools for Visual Storytelling

2026 Latest Best Top-Rated AI Dialogue Bubble Tools for Visual Storytelling are here on XIX.AI! This curated collection features powerful, game-changing tools that help creators boost productivity and overcome creative bottlenecks. Get a free vs paid comparison, see real-world tests, and check the latest rankings to find the must-try solutions perfect for crafting engaging visual narratives. Explore now to discover your ideal tool!

10 tools
xix.ai
Meeting Assistant Top AI Meeting Summary Tools: Track Decisions and Follow-Ups Clearly
Top AI Meeting Summary Tools: Track Decisions and Follow-Ups Clearly

2026 Latest Top-Rated Best AI Meeting Summary Tools for Clear Decision Tracking and Effortless Follow-Ups. This curated list showcases powerful, game-changing solutions that boost productivity dramatically by automating meeting notes, identifying key action items, and streamlining team coordination across all projects. Get a free vs paid comparison along with real-world tests and weekly updated rankings to help you find the perfect tool. Explore now to unlock your AI edge!

9 tools
xix.ai
Comments (8)
0/500
JoseRoberts
JoseRoberts August 12, 2025 at 11:00:59 AM EDT

This self-correction stuff is wild! 😮 It's like AI is learning to double-check its own homework. Wonder how far this 'agentic AI' will go—could it outsmart us at our own jobs soon?

WalterAnderson
WalterAnderson July 31, 2025 at 7:35:39 AM EDT

It's wild to think AI can now self-correct! 😮 Makes me wonder how soon we'll see these 'agentic AI' systems running our lives—hope they don’t outsmart us too much!

RonaldMartinez
RonaldMartinez July 22, 2025 at 3:39:52 AM EDT

This article really opened my eyes to how fast AI is evolving! Self-correcting LLMs sound like a game-changer for agentic AI. Can’t wait to see what new apps come out of this! 😄

WillieJackson
WillieJackson April 18, 2025 at 3:00:28 AM EDT

La perspectiva de Anthropic sobre que la IA no se estanca sino que supera los benchmarks es bastante genial. Es como si la IA estuviera jugando ajedrez mientras nosotros aún estamos tratando de entender las damas. Lo de la autocorrección suena prometedor, pero aún estoy un poco escéptico. 🤔

GeorgeWilson
GeorgeWilson April 17, 2025 at 1:45:24 PM EDT

Anthropic의 AI가 정체되지 않고 벤치마크를 뛰어넘는다는 생각이 멋지네요. AI는 체스를 하고 있는데, 우리는 아직 체커를 이해하는 단계예요. 자기 교정 이야기는 유망하지만, 아직 조금 회의적이에요. 🤔

NicholasCarter
NicholasCarter April 17, 2025 at 7:27:31 AM EDT

Anthropic's take on AI not stalling but outsmarting benchmarks is pretty cool. It's like AI is playing chess while we're still figuring out checkers. The self-correction stuff sounds promising, but I'm still a bit skeptical. 🤔

OR