option
Home
News
Washington State University Study Reveals ChatGPT's Major Contradictions in Complex Scientific Judgments

Washington State University Study Reveals ChatGPT's Major Contradictions in Complex Scientific Judgments

July 17, 2026
85

Washington State University Study Reveals ChatGPT

A recent study from Washington State University (WSU) reveals that while ChatGPT responds with confidence, its handling of complex scientific statements resembles random guessing. The research highlights not only limited accuracy but also frequent contradictory answers to the same question.

Professor Mesut Cicek and his team extracted 719 research hypotheses from business journals published since 2021 and repeatedly submitted them to the model for truth verification.

Although ChatGPT's surface accuracy appears around 80%, after accounting for random guessing, its actual performance is only about 60% better than a 50% coin flip. Researchers rated this as a "low D-grade score." The model performed especially poorly at identifying false statements, correctly judging only 16.4% of false propositions.

The researchers submitted each hypothesis to the model ten times and found it struggled to maintain consistent positions:

Answer Fluctuation: In about 73% of cases, the model maintained consistent conclusions across ten repetitions.

Extreme Contradictions: In some cases, the model alternated between "true" and "false" answers, with extreme instances where half were true and half false, despite using the exact same prompt.

The study notes that users are easily misled by AI's fluent and persuasive language, but this does not indicate genuine reasoning ability.

No True Understanding: The model relies on memory and pattern matching, unlike humans who genuinely understand the world and what they are saying.

Limited Version Progress: Testing showed that the updated ChatGPT‑5 mini (tested in 2025) performed similarly to earlier versions on this specific task, with no significant improvements.

Based on these findings, Cicek advises business managers to maintain a high degree of skepticism when making complex decisions. They should not regard generative AI as an "authority" that can replace professional judgment and must manually verify all outputs. Organizations should enhance training to help employees understand both the strengths and limitations of AI tools, preventing decision biases from blind trust.

This study serves as another reminder that despite rapid AI development, the technology's deep logical reasoning and evidence evaluation capabilities still require improvement.

Related article
Xiaopeng Humanoid Robot Factory in Guangzhou Starts Trial Production, Eyes 2026 for Mass Output Xiaopeng Humanoid Robot Factory in Guangzhou Starts Trial Production, Eyes 2026 for Mass Output Xiaopeng’s humanoid robot has entered small-scale trial production at its Guangzhou facility. The mass production line is currently undergoing final integration, signaling the start of the countdown to full-scale manufacturing.Previously, He Xiaopeng
Musk Calls for AI Giants to Cross-Test Models and Invite Rival Criticism Musk Calls for AI Giants to Cross-Test Models and Invite Rival Criticism During the All-In summit on September 15, Elon Musk, the world’s wealthiest individual, joined remotely via video. Addressing artificial intelligence, he urged top AI firms to cross-test each other’s models prior to public release, emphasizing that s
Claude Expands AI Office Coverage as Revision Mode Takes Center Stage in Ongoing Talks Claude Expands AI Office Coverage as Revision Mode Takes Center Stage in Ongoing Talks Anthropic has recently launched the public beta of Claude for Word, marking a significant milestone in its integration with Microsoft Office. Within just six months, Claude has embedded itself into the core trio of Office applications—Excel, PowerPoi
Related Special Topic Recommendations
Health & Wellness Top AI Fitness Coaches for Home Workouts: Train with Adaptive Plans
Top AI Fitness Coaches for Home Workouts: Train with Adaptive Plans

2026 Latest Top-rated AI Fitness Coaches for Home Workouts are here on XIX.AI! Our curated list features powerful, game-changing tools that create adaptive workout plans tailored to your goals, helping you boost productivity and break through fitness plateaus. Get a free vs paid comparison along with real-world tests and updated rankings weekly. Explore now to Discover your perfect tool and Start creating today to Unlock your AI edge!

10 tools
xix.ai
Social Media AI Comment Moderation Tools for TikTok Shops and High-Volume Brand Communities
AI Comment Moderation Tools for TikTok Shops and High-Volume Brand Communities

2026 Latest Best Top-rated AI Comment Moderation Tools for TikTok Shops and High-Volume Brand Communities are curated here by XIX.AI. Our weekly updated rankings feature powerful, game-changing solutions that deliver real-world tests results to help you boost content quality, enhance writing efficiency, and streamline community management effortlessly. Free vs paid comparison insights are also available. Explore now to Discover your perfect tool for unlocking your AI edge!

9 tools
xix.ai
Video creation CapCut AI Video Editors for Reels, Shorts, Ads, and Social Commerce Launches
CapCut AI Video Editors for Reels, Shorts, Ads, and Social Commerce Launches

2026 Latest Best CapCut AI Video Editors for Reels Shorts Ads and Social Commerce Launches! This curated list features top-rated powerful game-changing tools that deliver impressive results through real-world tests. You’ll find a free vs paid comparison along with detailed rankings to help you choose the perfect fit. XIX.AI offers expert insights to help you unlock your AI edge in content creation. Explore now to discover your ideal tool and boost your productivity today!

14 tools
xix.ai
Academic Research Best AI Study Discovery Tools: Find Journals, Datasets, and Trends Quickly
Best AI Study Discovery Tools: Find Journals, Datasets, and Trends Quickly

2026 Latest Best Top-rated AI Study Discovery Tools are here on XIX.AI! This curated list includes powerful game-changing resources for quick access to journals, datasets, and industry trends. Get a free vs paid comparison along with real-world tests and updated rankings weekly. Must-try options help boost your research efficiency and unlock your AI edge. Explore now to discover your perfect tool!

10 tools
xix.ai
Marketing AI Email Campaign Tools for Lifecycle Marketing
AI Email Campaign Tools for Lifecycle Marketing

2026 Latest Best Top-Rated AI Email Campaign Tools for Lifecycle Marketing! XIX.AI has curated a powerful, game-changing collection of must-try solutions that boost writing efficiency, help create highly targeted content, and streamline entire marketing workflows. Enjoy weekly updated rankings, real-world tests, and a free vs paid comparison to find the perfect tool for your needs. Explore now to unlock your AI edge!

11 tools
xix.ai
Finance Top AI Investment Research Tools: Compare Signals and Portfolio Risks
Top AI Investment Research Tools: Compare Signals and Portfolio Risks

2026 Latest Top-Rated Best AI Investment Research Tools are here on XIX.AI! This curated list includes powerfully accurate signals to assess portfolio risks and help you make smarter, data-driven decisions. We conduct real-world tests weekly to update rankings and offer a free vs paid comparison. Discover your perfect tool to boost investment analysis efficiency today. Explore now!

10 tools
xix.ai
Comments (1)
0/500
ScottKing
ScottKing September 27, 2026 at 4:00:08 PM EDT

WSUの研究結果、ChatGPTの科学的判断は実はランダムな推測に過ぎないって話か。自信満々に答えておきながら精度が低いのは困るな。医療や法律の分野で使おうとすると危険すぎる。AIの限界を再認識させられたよ。

OR