Home
OpenAI Slams AI Benchmark: Nearly a Third of Questions Flawed, Pass Rate Soared to 80% in 8 Months
OpenAI has publicly challenged the industry's leading benchmark, SWE-Bench Pro, in a blog post, claiming that roughly 30% of its 731 public test tasks contain evaluation flaws. Developed by Scale AI, SWE-Bench Pro is designed to assess the coding abilities of large language models and AI agents. Because it closely mirrors real-world enterprise development and enforces strict anti-cheating measures, it has become a widely trusted benchmark in AI software engineering.

OpenAI highlights a key signal in the blog post: the pass rate for top-tier models on this benchmark jumped from 23.3% to 80.3% in just eight months. Such rapid progress seems suspicious, and OpenAI argues that the benchmark can no longer accurately measure a model's real-world software development skills. The issue likely stems from flaws in the evaluation itself, not a genuine leap in model capability.
Two review paths cross-validated reveal that nearly 30% of tasks are deemed "unqualified."
To verify this, OpenAI launched two parallel review processes. The data-point analysis uncovered 200 failing tasks, or 27.4% of the 731 public tasks. Meanwhile, manual annotation flagged 249 failing tasks, or 34.1%. Cross-referencing both methods, OpenAI estimates that roughly 30% of SWE-Bench Pro tasks contain defects, falling into four categories: overly strict tests, inadequate prompts, narrow test coverage, and misleading prompts.
OpenAI also shared a typical example: one task asked for a single space at the start of a line when converting content to Markdown, but the hidden test expected two spaces. As a result, even if the model followed the stated requirement, it would be marked incorrect. This mismatch between explicit instructions and hidden requirements directly skews the evaluation of a model's true ability, and explains the unreasonable surge in pass rates.
Withdrawal of adoption recommendation and call for rebuilding the AI evaluation system
Based on this analysis, OpenAI has officially retracted its earlier recommendation to use SWE-Bench Pro. The company believes that future benchmarks should be created by experienced software developers specifically for AI evaluation, rather than repurposing test logic designed for human developers. If an industry benchmark has nearly 30% of its tasks flawed, the entire AI evaluation system's credibility comes into question. Shifting the focus from score-chasing competitions to genuine engineering capability assessments may be the next critical step in AI software engineering evaluation.
Related article
Baidu Q2 Revenue Hits 31.3 Billion Yuan as AI Business Surpasses Half for Second Straight Quarter
Baidu’s second-quarter 2026 financial results show total revenue of 31.3 billion yuan, with core business revenue at 25.2 billion yuan. AI-driven revenue now represents 50% of total business income, surpassing the halfway mark for two consecutive qua
OpenAI overtakes Anthropic in enterprise market share as AI spending grows
Recent figures from Ramp, a corporate credit card and expense management platform, reveal that OpenAI has reclaimed over 40% of the U.S. enterprise AI market, overtaking Anthropic. In May, Anthropic held a 41% share against OpenAI’s 39%; by July, Ant
Epson Unveils AX6 Cobot Featuring Compact Design and No-Code Programming
The AX6 force- and power-limited robot arm is designed to be lightweight and easy to use. Source: EpsonEpson Robots this week expanded its six-axis robot line with the AX6 collaborative robot. It said the new cobot is is designed with power- and forc
Related Special Topic Recommendations
Comments (0)
0/500
OpenAI has publicly challenged the industry's leading benchmark, SWE-Bench Pro, in a blog post, claiming that roughly 30% of its 731 public test tasks contain evaluation flaws. Developed by Scale AI, SWE-Bench Pro is designed to assess the coding abilities of large language models and AI agents. Because it closely mirrors real-world enterprise development and enforces strict anti-cheating measures, it has become a widely trusted benchmark in AI software engineering.

OpenAI highlights a key signal in the blog post: the pass rate for top-tier models on this benchmark jumped from 23.3% to 80.3% in just eight months. Such rapid progress seems suspicious, and OpenAI argues that the benchmark can no longer accurately measure a model's real-world software development skills. The issue likely stems from flaws in the evaluation itself, not a genuine leap in model capability.
Two review paths cross-validated reveal that nearly 30% of tasks are deemed "unqualified."
To verify this, OpenAI launched two parallel review processes. The data-point analysis uncovered 200 failing tasks, or 27.4% of the 731 public tasks. Meanwhile, manual annotation flagged 249 failing tasks, or 34.1%. Cross-referencing both methods, OpenAI estimates that roughly 30% of SWE-Bench Pro tasks contain defects, falling into four categories: overly strict tests, inadequate prompts, narrow test coverage, and misleading prompts.
OpenAI also shared a typical example: one task asked for a single space at the start of a line when converting content to Markdown, but the hidden test expected two spaces. As a result, even if the model followed the stated requirement, it would be marked incorrect. This mismatch between explicit instructions and hidden requirements directly skews the evaluation of a model's true ability, and explains the unreasonable surge in pass rates.
Withdrawal of adoption recommendation and call for rebuilding the AI evaluation system
Based on this analysis, OpenAI has officially retracted its earlier recommendation to use SWE-Bench Pro. The company believes that future benchmarks should be created by experienced software developers specifically for AI evaluation, rather than repurposing test logic designed for human developers. If an industry benchmark has nearly 30% of its tasks flawed, the entire AI evaluation system's credibility comes into question. Shifting the focus from score-chasing competitions to genuine engineering capability assessments may be the next critical step in AI software engineering evaluation.
Baidu Q2 Revenue Hits 31.3 Billion Yuan as AI Business Surpasses Half for Second Straight Quarter
Baidu’s second-quarter 2026 financial results show total revenue of 31.3 billion yuan, with core business revenue at 25.2 billion yuan. AI-driven revenue now represents 50% of total business income, surpassing the halfway mark for two consecutive qua
OpenAI overtakes Anthropic in enterprise market share as AI spending grows
Recent figures from Ramp, a corporate credit card and expense management platform, reveal that OpenAI has reclaimed over 40% of the U.S. enterprise AI market, overtaking Anthropic. In May, Anthropic held a 41% share against OpenAI’s 39%; by July, Ant
Epson Unveils AX6 Cobot Featuring Compact Design and No-Code Programming
The AX6 force- and power-limited robot arm is designed to be lightweight and easy to use. Source: EpsonEpson Robots this week expanded its six-axis robot line with the AX6 collaborative robot. It said the new cobot is is designed with power- and forc











