Claude Opus 4.7 Launches with Reliability Valued Over Intelligence
Anthropic has maintained an aggressive pace this year, rolling out new features almost every other day. The much-anticipated Claude Opus 4.7 has just been officially released, and interestingly, Anthropic was upfront in the announcement: "This is not our most powerful model." The rumored, stronger Claude Mythos Preview remains on standby. Still, Opus 4.7 has generated considerable attention because it tackles the issue of being "more reliable" rather than "smarter."

Benchmark results are notably impressive. On the rigorous coding benchmark SWE-bench Pro, 4.7 jumped from 53.4% in the previous version to 64.3%, a gain of nearly 11 percentage points, surpassing GPT-5.4 (57.7%) and Gemini 3.1 Pro (54.2%). On the visual reasoning benchmark CharXiv, it rose from 69.1% to 82.1%, driven by the newly added 2576-pixel long-side recognition capability, offering more than three times the clarity of its predecessor. On the tool call evaluation MCP-Atlas, it scored 77.3%, and on the legal AI platform Harvey's BigLaw benchmark, it reached 90.9%. However, on the agentic search evaluation BrowseComp, 4.7 saw a slight decline from 83.7% to 79.3%, overtaken by GPT-5.4 and Gemini—this is attributed to its "no fabrications" personality, preferring to report errors rather than guess when information is incomplete.
Beyond the numbers, the shift in temperament is more noteworthy. Replit's leader noted after testing: "It challenges me in technical discussions, helps me make better decisions, and truly acts like a better colleague." Data science platform Hex also observed that 4.7 directly reports errors when data is missing, rather than providing a "seemingly reasonable but completely incorrect" alternative value as before. At the same time, task resilience has improved significantly—Notion team tests indicate that the tool error rate has been reduced to one-third of previous levels, and when the tool chain fails, it can navigate obstacles and complete tasks independently. Vercel even discovered a new behavior: before writing system-level code, 4.7 first performs mathematical proofs on its own.

Of course, increased capability comes with a cost. 4.7 introduces a new tokenizer, generating 1 to 1.35 times more tokens for the same text. Additionally, it tends to "think a bit longer" on complex tasks, so actual consumption is almost certainly higher. To address this, Anthropic added an xhigh ultra-high thinking intensity level. Claude Code has set all packages to this level by default, and also launched the Deep Review instruction / ultrareview, Auto Mode extension for Max users, and a public beta version of the "task budget" feature to help developers manage token usage.
The more powerful Mythos Preview was recently made available to enterprises under the name "Project Glasswing" for cybersecurity research, but due to its overwhelming capability and incomplete security evaluations, it has not been publicly released yet.
Today's 4.7 represents the latest milestone in Anthropic's high-frequency delivery rhythm. Mythos will eventually arrive—and when it does, the already strong 4.7 may prove to be just the beginning.
Related article
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage
California AV Compliance: A New Era of Tickets, Geofences, and 1M Miles
Guident operates an AuveTech shuttle in South Florida, managing a four-mile route in West Palm Beach and a one-mile route in Boca Raton using its remote monitoring technology. | Credit: GuidentCalifornia is redefining the regulatory landscape for dri
Related Special Topic Recommendations
Comments (0)
0/500
Anthropic has maintained an aggressive pace this year, rolling out new features almost every other day. The much-anticipated Claude Opus 4.7 has just been officially released, and interestingly, Anthropic was upfront in the announcement: "This is not our most powerful model." The rumored, stronger Claude Mythos Preview remains on standby. Still, Opus 4.7 has generated considerable attention because it tackles the issue of being "more reliable" rather than "smarter."

Benchmark results are notably impressive. On the rigorous coding benchmark SWE-bench Pro, 4.7 jumped from 53.4% in the previous version to 64.3%, a gain of nearly 11 percentage points, surpassing GPT-5.4 (57.7%) and Gemini 3.1 Pro (54.2%). On the visual reasoning benchmark CharXiv, it rose from 69.1% to 82.1%, driven by the newly added 2576-pixel long-side recognition capability, offering more than three times the clarity of its predecessor. On the tool call evaluation MCP-Atlas, it scored 77.3%, and on the legal AI platform Harvey's BigLaw benchmark, it reached 90.9%. However, on the agentic search evaluation BrowseComp, 4.7 saw a slight decline from 83.7% to 79.3%, overtaken by GPT-5.4 and Gemini—this is attributed to its "no fabrications" personality, preferring to report errors rather than guess when information is incomplete.
Beyond the numbers, the shift in temperament is more noteworthy. Replit's leader noted after testing: "It challenges me in technical discussions, helps me make better decisions, and truly acts like a better colleague." Data science platform Hex also observed that 4.7 directly reports errors when data is missing, rather than providing a "seemingly reasonable but completely incorrect" alternative value as before. At the same time, task resilience has improved significantly—Notion team tests indicate that the tool error rate has been reduced to one-third of previous levels, and when the tool chain fails, it can navigate obstacles and complete tasks independently. Vercel even discovered a new behavior: before writing system-level code, 4.7 first performs mathematical proofs on its own.

Of course, increased capability comes with a cost. 4.7 introduces a new tokenizer, generating 1 to 1.35 times more tokens for the same text. Additionally, it tends to "think a bit longer" on complex tasks, so actual consumption is almost certainly higher. To address this, Anthropic added an xhigh ultra-high thinking intensity level. Claude Code has set all packages to this level by default, and also launched the Deep Review instruction / ultrareview, Auto Mode extension for Max users, and a public beta version of the "task budget" feature to help developers manage token usage.
The more powerful Mythos Preview was recently made available to enterprises under the name "Project Glasswing" for cybersecurity research, but due to its overwhelming capability and incomplete security evaluations, it has not been publicly released yet.
Today's 4.7 represents the latest milestone in Anthropic's high-frequency delivery rhythm. Mythos will eventually arrive—and when it does, the already strong 4.7 may prove to be just the beginning.
DeepMind CEO Hassabis: I sleep six hours a day, usually feel energetic around 1 a.m.
Fortune recently featured an interview with Demis Hassabis, CEO of Google DeepMind, revealing his unconventional approach to rest and productivity. Hassabis disclosed that he sleeps very little, structuring his waking hours into two distinct work blo
OpenAI, Anthropic Vie for Market Share Despite Revenue Shortfalls
Despite recent reports suggesting OpenAI missed revenue targets, creating pressure on tech stocks this Tuesday, private AI lab investors remain resilient. Seasoned backers have confirmed they will not reduce investment despite negative media coverage





Home






