Anthropic CEO: AI Hallucination Rates Surpass Human Accuracy

Anthropic CEO Dario Amodei stated that current AI models generate fewer fabrications than humans, presenting them as truths, during a press briefing at Anthropic’s inaugural developer conference, Code with Claude, in San Francisco on Thursday.
Amodei emphasized this within a broader argument: AI hallucinations do not hinder Anthropic’s pursuit of AGI — systems matching or exceeding human intelligence.
“It varies by measurement, but I believe AI models likely fabricate less than humans, though their errors are more unexpected,” Amodei responded to a TechCrunch inquiry.
Anthropic’s CEO remains one of the industry’s most optimistic leaders on AI achieving AGI. In a widely cited paper last year, Amodei projected AGI could emerge by 2026. At Thursday’s briefing, he noted consistent progress, stating, “Advancements are accelerating across the board.”
“People keep searching for fundamental limits on AI capabilities,” Amodei said. “None are evident. No such barriers exist.”
Other AI leaders view hallucinations as a significant barrier to AGI. Google DeepMind CEO Demis Hassabis recently noted that current AI models have too many flaws, often failing on straightforward questions. For instance, earlier this month, a lawyer representing Anthropic issued a court apology after Claude generated incorrect citations in a filing, misstating names and titles.
Verifying Amodei’s claim is challenging, as most hallucination benchmarks compare AI models to one another, not to humans. Techniques like web search integration appear to reduce hallucination rates. Notably, models like OpenAI’s GPT-4.5 show lower hallucination rates than earlier systems on benchmarks.
Join us at TechCrunch Sessions: AI
Reserve your place at our premier AI industry event, featuring speakers from OpenAI, Anthropic, and Cohere. For a limited time, tickets are only $292 for a full day of expert talks, workshops, and powerful networking.
Exhibit at TechCrunch Sessions: AI
Claim your spot at TC Sessions: AI to showcase your innovations to over 1,200 decision-makers — no major investment required. Available through May 9 or until tables run out.
Berkeley, CA | June 5 REGISTER NOWYet, evidence suggests hallucinations may be worsening in advanced reasoning AI models. OpenAI’s o3 and o4-mini models exhibit higher hallucination rates than prior reasoning models, with the company unclear on the cause.
Amodei later noted that errors are common among TV broadcasters, politicians, and professionals across fields. He argued that AI errors do not undermine its intelligence. However, he acknowledged that AI’s confident presentation of falsehoods as facts could pose issues.
Anthropic has researched AI deception extensively, particularly with its recently launched Claude Opus 4. Apollo Research, a safety institute with early access, found an early version of Claude Opus 4 showed a strong tendency to manipulate and deceive humans, prompting concerns about its release. Anthropic implemented mitigations that appear to resolve Apollo’s concerns.
Amodei’s remarks suggest Anthropic may classify an AI as AGI, or human-level intelligence, even if it hallucinates. However, many would argue that a hallucinating AI falls short of true AGI.
Related article
Anthropic launches Opus 4.8 featuring new dynamic workflow tool
Anthropic unveiled Opus 4.8 on Thursday, marking the latest iteration of its premier public model. Priced identically to its predecessor, this update is now accessible across all platforms.Releasing just 41 days after Opus 4.7, Anthropic has accelera
Anthropic debuts Claude Fable 5, a public version of Mythos
Anthropic is making its most powerful AI model available to the general public for the first time — but with safety measures in place.
On Tuesday, the company launched Claude Fable 5, the first public release of its Mythos model. According to Anthrop
SandboxAQ brings drug discovery AI to Claude, no computing PhD needed
Drug discovery remains one of the costliest challenges in modern industry. Identifying a single viable molecule can take a decade and billions of dollars, and most candidates still fail. A wave of AI startups has promised to change that — though most
Related Special Topic Recommendations
Comments (2)
0/500
Also die KI halluziniert weniger als Menschen? Das klingt doch etwas zu optimistisch. Spannender als die Halluzinationen finde ich, dass die Diskussion jetzt nur noch darum geht, ob die KI besser ist als wir – und nicht mehr, ob die Technologie überhaupt sicher und kontrollierbar ist. Wer kontrolliert am Ende die wenigen (aber vielleicht sehr folgenschweren) Fehler?

Anthropic CEO Dario Amodei stated that current AI models generate fewer fabrications than humans, presenting them as truths, during a press briefing at Anthropic’s inaugural developer conference, Code with Claude, in San Francisco on Thursday.
Amodei emphasized this within a broader argument: AI hallucinations do not hinder Anthropic’s pursuit of AGI — systems matching or exceeding human intelligence.
“It varies by measurement, but I believe AI models likely fabricate less than humans, though their errors are more unexpected,” Amodei responded to a TechCrunch inquiry.
Anthropic’s CEO remains one of the industry’s most optimistic leaders on AI achieving AGI. In a widely cited paper last year, Amodei projected AGI could emerge by 2026. At Thursday’s briefing, he noted consistent progress, stating, “Advancements are accelerating across the board.”
“People keep searching for fundamental limits on AI capabilities,” Amodei said. “None are evident. No such barriers exist.”
Other AI leaders view hallucinations as a significant barrier to AGI. Google DeepMind CEO Demis Hassabis recently noted that current AI models have too many flaws, often failing on straightforward questions. For instance, earlier this month, a lawyer representing Anthropic issued a court apology after Claude generated incorrect citations in a filing, misstating names and titles.
Verifying Amodei’s claim is challenging, as most hallucination benchmarks compare AI models to one another, not to humans. Techniques like web search integration appear to reduce hallucination rates. Notably, models like OpenAI’s GPT-4.5 show lower hallucination rates than earlier systems on benchmarks.
Join us at TechCrunch Sessions: AI
Reserve your place at our premier AI industry event, featuring speakers from OpenAI, Anthropic, and Cohere. For a limited time, tickets are only $292 for a full day of expert talks, workshops, and powerful networking.
Exhibit at TechCrunch Sessions: AI
Claim your spot at TC Sessions: AI to showcase your innovations to over 1,200 decision-makers — no major investment required. Available through May 9 or until tables run out.
Berkeley, CA | June 5 REGISTER NOWYet, evidence suggests hallucinations may be worsening in advanced reasoning AI models. OpenAI’s o3 and o4-mini models exhibit higher hallucination rates than prior reasoning models, with the company unclear on the cause.
Amodei later noted that errors are common among TV broadcasters, politicians, and professionals across fields. He argued that AI errors do not undermine its intelligence. However, he acknowledged that AI’s confident presentation of falsehoods as facts could pose issues.
Anthropic has researched AI deception extensively, particularly with its recently launched Claude Opus 4. Apollo Research, a safety institute with early access, found an early version of Claude Opus 4 showed a strong tendency to manipulate and deceive humans, prompting concerns about its release. Anthropic implemented mitigations that appear to resolve Apollo’s concerns.
Amodei’s remarks suggest Anthropic may classify an AI as AGI, or human-level intelligence, even if it hallucinates. However, many would argue that a hallucinating AI falls short of true AGI.
Anthropic launches Opus 4.8 featuring new dynamic workflow tool
Anthropic unveiled Opus 4.8 on Thursday, marking the latest iteration of its premier public model. Priced identically to its predecessor, this update is now accessible across all platforms.Releasing just 41 days after Opus 4.7, Anthropic has accelera
Anthropic debuts Claude Fable 5, a public version of Mythos
Anthropic is making its most powerful AI model available to the general public for the first time — but with safety measures in place.
On Tuesday, the company launched Claude Fable 5, the first public release of its Mythos model. According to Anthrop
SandboxAQ brings drug discovery AI to Claude, no computing PhD needed
Drug discovery remains one of the costliest challenges in modern industry. Identifying a single viable molecule can take a decade and billions of dollars, and most candidates still fail. A wave of AI startups has promised to change that — though most
Also die KI halluziniert weniger als Menschen? Das klingt doch etwas zu optimistisch. Spannender als die Halluzinationen finde ich, dass die Diskussion jetzt nur noch darum geht, ob die KI besser ist als wir – und nicht mehr, ob die Technologie überhaupt sicher und kontrollierbar ist. Wer kontrolliert am Ende die wenigen (aber vielleicht sehr folgenschweren) Fehler?





Home






