Home
Washington State University Study Reveals ChatGPT's Major Contradictions in Complex Scientific Judgments

A recent study from Washington State University (WSU) reveals that while ChatGPT responds with confidence, its handling of complex scientific statements resembles random guessing. The research highlights not only limited accuracy but also frequent contradictory answers to the same question.
Professor Mesut Cicek and his team extracted 719 research hypotheses from business journals published since 2021 and repeatedly submitted them to the model for truth verification.
Although ChatGPT's surface accuracy appears around 80%, after accounting for random guessing, its actual performance is only about 60% better than a 50% coin flip. Researchers rated this as a "low D-grade score." The model performed especially poorly at identifying false statements, correctly judging only 16.4% of false propositions.
The researchers submitted each hypothesis to the model ten times and found it struggled to maintain consistent positions:
Answer Fluctuation: In about 73% of cases, the model maintained consistent conclusions across ten repetitions.
Extreme Contradictions: In some cases, the model alternated between "true" and "false" answers, with extreme instances where half were true and half false, despite using the exact same prompt.
The study notes that users are easily misled by AI's fluent and persuasive language, but this does not indicate genuine reasoning ability.
No True Understanding: The model relies on memory and pattern matching, unlike humans who genuinely understand the world and what they are saying.
Limited Version Progress: Testing showed that the updated ChatGPT‑5 mini (tested in 2025) performed similarly to earlier versions on this specific task, with no significant improvements.
Based on these findings, Cicek advises business managers to maintain a high degree of skepticism when making complex decisions. They should not regard generative AI as an "authority" that can replace professional judgment and must manually verify all outputs. Organizations should enhance training to help employees understand both the strengths and limitations of AI tools, preventing decision biases from blind trust.
This study serves as another reminder that despite rapid AI development, the technology's deep logical reasoning and evidence evaluation capabilities still require improvement.
Related article
Xiaopeng Humanoid Robot Factory in Guangzhou Starts Trial Production, Eyes 2026 for Mass Output
Xiaopeng’s humanoid robot has entered small-scale trial production at its Guangzhou facility. The mass production line is currently undergoing final integration, signaling the start of the countdown to full-scale manufacturing.Previously, He Xiaopeng
Musk Calls for AI Giants to Cross-Test Models and Invite Rival Criticism
During the All-In summit on September 15, Elon Musk, the world’s wealthiest individual, joined remotely via video. Addressing artificial intelligence, he urged top AI firms to cross-test each other’s models prior to public release, emphasizing that s
Claude Expands AI Office Coverage as Revision Mode Takes Center Stage in Ongoing Talks
Anthropic has recently launched the public beta of Claude for Word, marking a significant milestone in its integration with Microsoft Office. Within just six months, Claude has embedded itself into the core trio of Office applications—Excel, PowerPoi
Related Special Topic Recommendations
Comments (1)
0/500

A recent study from Washington State University (WSU) reveals that while ChatGPT responds with confidence, its handling of complex scientific statements resembles random guessing. The research highlights not only limited accuracy but also frequent contradictory answers to the same question.
Professor Mesut Cicek and his team extracted 719 research hypotheses from business journals published since 2021 and repeatedly submitted them to the model for truth verification.
Although ChatGPT's surface accuracy appears around 80%, after accounting for random guessing, its actual performance is only about 60% better than a 50% coin flip. Researchers rated this as a "low D-grade score." The model performed especially poorly at identifying false statements, correctly judging only 16.4% of false propositions.
The researchers submitted each hypothesis to the model ten times and found it struggled to maintain consistent positions:
Answer Fluctuation: In about 73% of cases, the model maintained consistent conclusions across ten repetitions.
Extreme Contradictions: In some cases, the model alternated between "true" and "false" answers, with extreme instances where half were true and half false, despite using the exact same prompt.
The study notes that users are easily misled by AI's fluent and persuasive language, but this does not indicate genuine reasoning ability.
No True Understanding: The model relies on memory and pattern matching, unlike humans who genuinely understand the world and what they are saying.
Limited Version Progress: Testing showed that the updated ChatGPT‑5 mini (tested in 2025) performed similarly to earlier versions on this specific task, with no significant improvements.
Based on these findings, Cicek advises business managers to maintain a high degree of skepticism when making complex decisions. They should not regard generative AI as an "authority" that can replace professional judgment and must manually verify all outputs. Organizations should enhance training to help employees understand both the strengths and limitations of AI tools, preventing decision biases from blind trust.
This study serves as another reminder that despite rapid AI development, the technology's deep logical reasoning and evidence evaluation capabilities still require improvement.
Xiaopeng Humanoid Robot Factory in Guangzhou Starts Trial Production, Eyes 2026 for Mass Output
Xiaopeng’s humanoid robot has entered small-scale trial production at its Guangzhou facility. The mass production line is currently undergoing final integration, signaling the start of the countdown to full-scale manufacturing.Previously, He Xiaopeng
Musk Calls for AI Giants to Cross-Test Models and Invite Rival Criticism
During the All-In summit on September 15, Elon Musk, the world’s wealthiest individual, joined remotely via video. Addressing artificial intelligence, he urged top AI firms to cross-test each other’s models prior to public release, emphasizing that s
Claude Expands AI Office Coverage as Revision Mode Takes Center Stage in Ongoing Talks
Anthropic has recently launched the public beta of Claude for Word, marking a significant milestone in its integration with Microsoft Office. Within just six months, Claude has embedded itself into the core trio of Office applications—Excel, PowerPoi











