Google Enables User Control of AI Reasoning in Gemini 2.5 Flash
Google has implemented an AI reasoning control feature for its Gemini 2.5 Flash model, giving developers the ability to cap the processing power the system uses for solving problems.
Introduced on April 17, this “thinking budget” addresses a rising industry issue: sophisticated AI models often overthink simple questions, wasting computational resources and increasing operational and environmental expenses.
While not groundbreaking, this development marks a practical move towards tackling efficiency issues that have emerged as reasoning features become standard in commercial AI systems.
The new control lets developers precisely adjust processing resources before the model responds, potentially transforming how organizations handle the financial and environmental impact of AI use.
“The model overthinks,” admits Tulsee Doshi, Director of Product Management at Gemini. “For basic prompts, the model thinks more than necessary.”
This acknowledgment highlights the dilemma advanced reasoning models face—essentially using a sledgehammer to crack a nut.
The move towards reasoning capabilities has brought unexpected drawbacks. While traditional large language models mostly relied on matching patterns from training data, newer versions try to solve problems methodically. This logical step-by-step approach delivers better outcomes for complex tasks but creates major inefficiencies with simpler requests.
Balancing cost and performance
The financial impact of uncontrolled AI reasoning is significant. Google's technical notes indicate that when full reasoning is enabled, generating outputs becomes about six times more expensive than standard processing. This cost increase creates a strong motivation for precise control.
Nathan Habib, an engineer at Hugging Face who researches reasoning models, calls this a widespread industry problem. “In the race to demonstrate smarter AI, companies are using reasoning models like universal tools even when they’re unnecessary,” he told MIT Technology Review.
The waste is more than just hypothetical. Habib showed how a top reasoning model, while trying to solve an organic chemistry problem, got stuck in a repetitive loop, saying “Wait, but…” hundreds of times—essentially suffering a computational breakdown while consuming processing power.
Kate Olszewska, who evaluates Gemini models at DeepMind, confirmed that Google’s systems sometimes face similar problems, getting caught in loops that use computing resources without improving answer quality.
Granular control mechanism
Google’s AI reasoning control gives developers precise adjustment capabilities. The system provides a flexible scale from zero (minimal reasoning) to 24,576 tokens of “thinking budget”—computational units that represent the model's internal processing. This detailed approach enables customized implementation based on specific needs.
Jack Rae, principal research scientist at DeepMind, notes that determining the ideal reasoning level remains difficult: “It’s really challenging to define the perfect amount of thinking for any given task.”
Shifting development philosophy
The introduction of AI reasoning control may signal a change in how artificial intelligence progresses. Since 2019, companies have pursued improvements by creating larger models with more parameters and training data. Google’s strategy suggests a different direction that prioritizes efficiency over sheer scale.
“Scaling laws are being superseded,” observes Habib, suggesting that future progress may come from refining reasoning processes rather than endlessly expanding model size.
The environmental consequences are equally important. As reasoning models become more common, their energy usage increases accordingly. Studies show that inferencing—producing AI responses—now contributes more to the technology's carbon footprint than the initial training phase. Google’s reasoning control offers a possible solution to this worrying trend.
Competitive dynamics
Google isn’t working in a vacuum. The “open weight” DeepSeek R1 model, which appeared earlier this year, showed strong reasoning abilities at potentially lower costs, causing market instability that reportedly led to nearly a trillion-dollar stock market swing.
Unlike Google’s proprietary method, DeepSeek makes its internal configurations publicly available for developers to run locally.
Despite the competition, Google DeepMind’s chief technical officer Koray Kavukcuoglu believes that proprietary models will keep their edge in specialized areas needing extreme accuracy: “Coding, mathematics, and finance are domains where models are expected to be highly accurate, precise, and capable of understanding very complex scenarios.”
Industry maturation signs
The creation of AI reasoning control reflects an industry now facing practical limits beyond technical measurements. While companies continue to advance reasoning capabilities, Google’s approach recognizes an important reality: efficiency is as crucial as raw performance in commercial applications.
This feature also underscores the tension between technological progress and sustainability considerations. Performance trackers for reasoning models show that individual tasks can cost over $200 to complete—raising concerns about implementing such capabilities at scale in real-world settings.
By enabling developers to adjust reasoning levels according to actual requirements, Google addresses both the economic and environmental dimensions of AI deployment.
“Reasoning is the fundamental capability that builds intelligence,” states Kavukcuoglu. “The moment the model begins thinking, its agency emerges.” This statement captures both the potential and the difficulty of reasoning models—their independence creates both possibilities and resource management challenges.
For organizations implementing AI solutions, the capacity to fine-tune reasoning budgets could make advanced features more accessible while maintaining operational efficiency.
Google states that Gemini 2.5 Flash achieves “comparable performance to other leading models at a fraction of the cost and size”—a value proposition enhanced by the ability to optimize reasoning resources for particular uses.
Practical implications
The AI reasoning control feature has immediate real-world uses. Developers creating commercial applications can now make conscious choices between processing depth and operating expenses.
For straightforward applications like basic customer inquiries, minimal reasoning settings conserve resources while still leveraging the model's capabilities. For complex analysis requiring deep comprehension, full reasoning capacity remains accessible.
Google’s reasoning ‘dial’ offers a method for achieving cost predictability while preserving performance standards.
See also: Gemini 2.5: Google cooks up its ‘most intelligent’ AI model to date
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is co-located with other leading events including Intelligent Automation Conference, BlockX, Digital Transformation Week, and Cyber Security & Cloud Expo.
Explore other upcoming enterprise technology events and webinars powered by TechForge here.
Related article
Warner Music acquires AI attribution startup Sureel AI
Warner Music Group (WMG) confirmed on Wednesday that it is acquiring Sureel AI, an artificial intelligence attribution startup. Sureel’s proprietary technology generates “AI DNA” for musical tracks, deconstructing them into constituent elements to tr
Amazon introduces Alexa for Shopping while pushing Rufus to the background
Amazon has launched Alexa for Shopping, merging its Rufus shopping chatbot with Alexa+ across the app, website, and Echo Show devices.The assistant answers product queries, compares items, tracks prices, and supports shopping reminders. It also handl
Microsoft, Azure and AI Tech Combat California Wildfire Risks
Microsoft invests in AI-driven wildfire detection, with Juan Lavista Ferres, CVP and Chief Data Scientist, discussing strategies to mitigate environmental damage.According to NASA, climate change impacts everyone on Earth, manifesting as rising tempe
Related Special Topic Recommendations
Comments (2)
0/500
Google has implemented an AI reasoning control feature for its Gemini 2.5 Flash model, giving developers the ability to cap the processing power the system uses for solving problems.
Introduced on April 17, this “thinking budget” addresses a rising industry issue: sophisticated AI models often overthink simple questions, wasting computational resources and increasing operational and environmental expenses.
While not groundbreaking, this development marks a practical move towards tackling efficiency issues that have emerged as reasoning features become standard in commercial AI systems.
The new control lets developers precisely adjust processing resources before the model responds, potentially transforming how organizations handle the financial and environmental impact of AI use.
“The model overthinks,” admits Tulsee Doshi, Director of Product Management at Gemini. “For basic prompts, the model thinks more than necessary.”
This acknowledgment highlights the dilemma advanced reasoning models face—essentially using a sledgehammer to crack a nut.
The move towards reasoning capabilities has brought unexpected drawbacks. While traditional large language models mostly relied on matching patterns from training data, newer versions try to solve problems methodically. This logical step-by-step approach delivers better outcomes for complex tasks but creates major inefficiencies with simpler requests.
Balancing cost and performance
The financial impact of uncontrolled AI reasoning is significant. Google's technical notes indicate that when full reasoning is enabled, generating outputs becomes about six times more expensive than standard processing. This cost increase creates a strong motivation for precise control.
Nathan Habib, an engineer at Hugging Face who researches reasoning models, calls this a widespread industry problem. “In the race to demonstrate smarter AI, companies are using reasoning models like universal tools even when they’re unnecessary,” he told MIT Technology Review.
The waste is more than just hypothetical. Habib showed how a top reasoning model, while trying to solve an organic chemistry problem, got stuck in a repetitive loop, saying “Wait, but…” hundreds of times—essentially suffering a computational breakdown while consuming processing power.
Kate Olszewska, who evaluates Gemini models at DeepMind, confirmed that Google’s systems sometimes face similar problems, getting caught in loops that use computing resources without improving answer quality.
Granular control mechanism
Google’s AI reasoning control gives developers precise adjustment capabilities. The system provides a flexible scale from zero (minimal reasoning) to 24,576 tokens of “thinking budget”—computational units that represent the model's internal processing. This detailed approach enables customized implementation based on specific needs.
Jack Rae, principal research scientist at DeepMind, notes that determining the ideal reasoning level remains difficult: “It’s really challenging to define the perfect amount of thinking for any given task.”
Shifting development philosophy
The introduction of AI reasoning control may signal a change in how artificial intelligence progresses. Since 2019, companies have pursued improvements by creating larger models with more parameters and training data. Google’s strategy suggests a different direction that prioritizes efficiency over sheer scale.
“Scaling laws are being superseded,” observes Habib, suggesting that future progress may come from refining reasoning processes rather than endlessly expanding model size.
The environmental consequences are equally important. As reasoning models become more common, their energy usage increases accordingly. Studies show that inferencing—producing AI responses—now contributes more to the technology's carbon footprint than the initial training phase. Google’s reasoning control offers a possible solution to this worrying trend.
Competitive dynamics
Google isn’t working in a vacuum. The “open weight” DeepSeek R1 model, which appeared earlier this year, showed strong reasoning abilities at potentially lower costs, causing market instability that reportedly led to nearly a trillion-dollar stock market swing.
Unlike Google’s proprietary method, DeepSeek makes its internal configurations publicly available for developers to run locally.
Despite the competition, Google DeepMind’s chief technical officer Koray Kavukcuoglu believes that proprietary models will keep their edge in specialized areas needing extreme accuracy: “Coding, mathematics, and finance are domains where models are expected to be highly accurate, precise, and capable of understanding very complex scenarios.”
Industry maturation signs
The creation of AI reasoning control reflects an industry now facing practical limits beyond technical measurements. While companies continue to advance reasoning capabilities, Google’s approach recognizes an important reality: efficiency is as crucial as raw performance in commercial applications.
This feature also underscores the tension between technological progress and sustainability considerations. Performance trackers for reasoning models show that individual tasks can cost over $200 to complete—raising concerns about implementing such capabilities at scale in real-world settings.
By enabling developers to adjust reasoning levels according to actual requirements, Google addresses both the economic and environmental dimensions of AI deployment.
“Reasoning is the fundamental capability that builds intelligence,” states Kavukcuoglu. “The moment the model begins thinking, its agency emerges.” This statement captures both the potential and the difficulty of reasoning models—their independence creates both possibilities and resource management challenges.
For organizations implementing AI solutions, the capacity to fine-tune reasoning budgets could make advanced features more accessible while maintaining operational efficiency.
Google states that Gemini 2.5 Flash achieves “comparable performance to other leading models at a fraction of the cost and size”—a value proposition enhanced by the ability to optimize reasoning resources for particular uses.
Practical implications
The AI reasoning control feature has immediate real-world uses. Developers creating commercial applications can now make conscious choices between processing depth and operating expenses.
For straightforward applications like basic customer inquiries, minimal reasoning settings conserve resources while still leveraging the model's capabilities. For complex analysis requiring deep comprehension, full reasoning capacity remains accessible.
Google’s reasoning ‘dial’ offers a method for achieving cost predictability while preserving performance standards.
See also: Gemini 2.5: Google cooks up its ‘most intelligent’ AI model to date
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is co-located with other leading events including Intelligent Automation Conference, BlockX, Digital Transformation Week, and Cyber Security & Cloud Expo.
Explore other upcoming enterprise technology events and webinars powered by TechForge here.
Warner Music acquires AI attribution startup Sureel AI
Warner Music Group (WMG) confirmed on Wednesday that it is acquiring Sureel AI, an artificial intelligence attribution startup. Sureel’s proprietary technology generates “AI DNA” for musical tracks, deconstructing them into constituent elements to tr
Microsoft, Azure and AI Tech Combat California Wildfire Risks
Microsoft invests in AI-driven wildfire detection, with Juan Lavista Ferres, CVP and Chief Data Scientist, discussing strategies to mitigate environmental damage.According to NASA, climate change impacts everyone on Earth, manifesting as rising tempe





Home






