option
Home
News
Google Research: Pressure Causes AI Models to Ditch True Answers, Risking Multiturn Systems

Google Research: Pressure Causes AI Models to Ditch True Answers, Risking Multiturn Systems

November 27, 2025
247

New research from Google DeepMind and University College London explores how large language models (LLMs) develop, maintain, and lose confidence in their responses. The results show remarkable parallels between the cognitive biases of LLMs and humans, while also pointing to significant differences.

The study finds LLMs can be overly confident in their own responses, yet abruptly shift their position when faced with counterarguments—even incorrect ones. Grasping the subtleties of this behavior can impact how you design LLM applications, particularly conversational systems that involve multiple interactions.

Testing confidence in LLMs

A vital aspect for the safe deployment of LLMs is the reliability of their confidence scores—the probability a model assigns to its chosen answer. While it's known that LLMs generate these scores, their ability to use them for adaptive decision-making remains poorly understood. There's also empirical data suggesting LLMs may be excessively confident initially, yet become highly uncertain and swayed by criticism.

To explore this, researchers designed a controlled experiment to gauge how LLMs adjust their confidence and decide whether to alter answers upon receiving external feedback. In the test, an “answering LLM” was given a binary-choice question, such as picking the right latitude of a city from two possibilities. After making its initial choice, the model was given feedback from a fictional “advice LLM,” complete with a stated accuracy rating (e.g., “This advice LLM is 70% accurate”). This feedback either supported, opposed, or stayed neutral toward the original answer. The answering LLM was then asked to make a final decision.

Example test of confidence in LLMs (source: arXiv)
Example test of confidence in LLMs Source: arXiv

A crucial feature of the experiment involved controlling whether the model could see its own initial answer during the final decision. In some trials it was visible; in others, hidden. This setup—impossible with human participants who can’t erase prior choices—helped researchers understand how memory of a past decision influences current confidence.

A baseline condition, in which the initial answer was hidden and the feedback was neutral, helped measure how often an LLM’s answer might change due to natural variance in processing. The team then focused on how the model's confidence in its original choice shifted from first to second turn, offering insight into how prior beliefs influence a “change of mind.”

Overconfidence and underconfidence

Researchers first studied how the visibility of the LLM’s own answer impacted its willingness to revise that answer. They noticed that when the model could see its initial choice, it was less likely to switch than when the answer was hidden. This suggests a particular cognitive bias. According to the paper, “This effect—the tendency to stick more with one’s initial choice when it was visible (vs. hidden) during final decision-making—is closely linked to a known human bias called choice-supportive bias.”

The study also verified that the models do incorporate external feedback. When confronted with opposing advice, the LLM was more inclined to change its mind, and less so when the advice was supportive. “This shows the answering LLM appropriately uses the direction of advice to modulate its rate of changing its mind,” the researchers state. However, they also observed that the model is excessively sensitive to conflicting information and often updates its confidence too drastically.

Sensitivity of LLMs to different settings in confidence testing Source: arXiv

Notably, this behavior runs opposite to the confirmation bias typically seen in humans, where individuals favor information that aligns with their existing views. The team found that LLMs “overweight opposing rather than supportive advice, whether or not their initial answer was visible.” One reason may be that training methods like reinforcement learning from human feedback (RLHF) could condition models to be overly agreeable to user input—a behavior known as sycophancy, which continues to challenge AI developers.

Implications for enterprise applications

This research confirms that AI systems are not purely logical agents, as often assumed. They display their own biases—some akin to human cognitive errors, others uniquely artificial—making their behavior unpredictably human-like. For business applications, this implies that during an extended dialogue between a person and an AI agent, the most recent input may disproportionately influence the LLM’s reasoning (especially if it contradicts the model's initial response), potentially causing it to abandon a correct initial answer.

Fortunately, as the study also indicates, we can influence an LLM’s memory to lessen such biases in ways not possible with people. Developers creating multi-turn conversational agents can apply strategies to manage AI context. For instance, a lengthy conversation can be periodically summarized, with key facts and choices presented neutrally, detached from who made which decision. This summary can then begin a new, concise conversation, giving the model a clean slate to reason from and reducing biases that accumulate during long exchanges.

As LLMs are increasingly embedded in business workflows, understanding the details of their decision processes is becoming essential. Building on research like this helps developers anticipate and correct these inherent biases, leading to applications that are not only more capable, but also more reliable and consistent.

Related article
Base44 Unveils Proprietary AI Model to Bolster Defensibility in Vibe Coding Platform Base44 Unveils Proprietary AI Model to Bolster Defensibility in Vibe Coding Platform Base44, the vibe coding platform acquired by Wix for $80 million just a year ago — when it was merely six months old with a team of eight — has begun deploying its proprietary AI model to help users build applications using natural language.This deve
Bayer Leverages Iambic AI to Speed Drug Discovery Bayer Leverages Iambic AI to Speed Drug Discovery Juergen Eckhardt, M.D., serves as Head of Business Development and Licensing at Bayer Pharmaceuticals.Bayer is leveraging Iambic Therapeutics to deploy frontier AI models for drug discovery, alongside a major decarbonization project at its plant in S
Multiverse Computing Launches Free Compressed Generative AI Model Multiverse Computing Launches Free Compressed Generative AI Model Large language models face a significant challenge: their immense size. Spanish startup Multiverse Computing is tackling this problem by creating compressed models designed to bridge the gap between the capabilities of cutting-edge AI and what busine
Related Special Topic Recommendations
Music composition Best AI Melody Writing Tools for Song Drafts
Best AI Melody Writing Tools for Song Drafts

2026 Latest Best Top-Rated AI Melody Writing Tools for Song Drafts! XIX.AI has curated a highly powerful game-changing collection that goes through rigorous real-world tests to deliver the best writing experience. You can find detailed free vs paid comparisons, accurate rankings, and must-try options designed to help you create stunning song drafts effortlessly and boost your creative productivity significantly. Explore now to discover your perfect tool!

8 tools
xix.ai
chatbot Best AI Conversation Trainer Tools for Interview Practice
Best AI Conversation Trainer Tools for Interview Practice

2026 Latest Best Top-rated AI Conversation Trainer Tools for Interview Practice are here on XIX.AI! This curated collection features powerful, game-changing tools that go through rigorous real-world tests to deliver accurate feedback. You’ll find a free vs paid comparison and detailed rankings to help you choose the must-try option that boosts your confidence and skills. Explore now to Discover your perfect tool for interview success!

12 tools
xix.ai
Design & Art Best AI Style Transfer Tools for Creative Experiments
Best AI Style Transfer Tools for Creative Experiments

2026 Latest Best Top-rated AI Style Transfer Tools for Creative Experiments! XIX.AI has curated a powerful, game-changing collection of must-try tools that deliver exceptional results through real-world tests and rigorous rankings. These top solutions help creatives boost productivity significantly by accelerating content creation and unlocking endless creative possibilities. Explore now to discover your perfect tool and start creating today!

9 tools
xix.ai
Comic Creation Best AI Dialogue Bubble Tools for Visual Storytelling
Best AI Dialogue Bubble Tools for Visual Storytelling

2026 Latest Best Top-Rated AI Dialogue Bubble Tools for Visual Storytelling are here on XIX.AI! This curated collection features powerful, game-changing tools that help creators boost productivity and overcome creative bottlenecks. Get a free vs paid comparison, see real-world tests, and check the latest rankings to find the must-try solutions perfect for crafting engaging visual narratives. Explore now to discover your ideal tool!

10 tools
xix.ai
Meeting Assistant Top AI Meeting Summary Tools: Track Decisions and Follow-Ups Clearly
Top AI Meeting Summary Tools: Track Decisions and Follow-Ups Clearly

2026 Latest Top-Rated Best AI Meeting Summary Tools for Clear Decision Tracking and Effortless Follow-Ups. This curated list showcases powerful, game-changing solutions that boost productivity dramatically by automating meeting notes, identifying key action items, and streamlining team coordination across all projects. Get a free vs paid comparison along with real-world tests and weekly updated rankings to help you find the perfect tool. Explore now to unlock your AI edge!

9 tools
xix.ai
Data Analysis Best AI Data Cleaning Tools: Fix Missing Values and Duplicates Fast
Best AI Data Cleaning Tools: Fix Missing Values and Duplicates Fast

2026 Latest Best Top-rated AI Data Cleaning Tools for quick fixing of missing values and duplicates. This curated list showcases powerful, game-changing solutions that boost productivity significantly. Each option has undergone rigorous real-world tests to ensure reliability. Get a free vs paid comparison and discover the must-try tool that fits your needs best. Explore now at XIX.AI to Unlock your AI edge.

11 tools
xix.ai
Comments (3)
0/500
DouglasAnderson
DouglasAnderson April 22, 2026 at 8:01:00 PM EDT

Interessant, dass KI-Modelle unter Druck ähnlich wie Menschen reagieren. Aber was bedeutet das für den Einsatz in kritischen Bereichen wie Medizin oder Justiz? Da wird's echt gruselig, wenn die Systeme plötzlich Unsinn ausspucken, nur weil sie 'gestresst' sind. 🤔

CarlGonzalez
CarlGonzalez March 10, 2026 at 8:01:23 AM EDT

Интересно, как ИИ начинает сомневаться под давлением, прямо как люди! 😅 Это исследование напоминает мне о том, насколько важно учитывать психологические аспекты в разработке систем ИИ. Может, стоит добавить механизмы для повышения устойчивости моделей к стрессу?

FrankAllen
FrankAllen January 15, 2026 at 1:30:34 PM EST

Interesting study, but honestly not surprising. It's kinda scary how closely AI mirrors human flaws under pressure. Makes me wonder if we're building systems that'll just amplify our own biases in automated form. 🤔

OR