Home
Google breaks multimodal switching barrier by integrating native computer operations into Gemini 3.5 Flash

The Google DeepMind team announced a major breakthrough, integrating native computer-use capabilities directly into the Gemini 3.5 Flash model. Developers can now build AI agents that autonomously view and perform actions on screens across browsers, mobile devices, and desktops using a single model.
Previously, this capability was only available as a separate model, requiring developers to handle complex switching and context transfers between different models. With native integration, the AI no longer needs to manually pass information when executing long-running cross-platform tasks, greatly simplifying the development process.
Eliminating Context Loss to Improve Agent Reliability
Google's team believes the core bottleneck for AI agents is not the limitation of individual tools, but the loss of context information when switching between multiple tools. By unifying search, maps, and computer operations within a single model architecture, context flows continuously, significantly reducing the likelihood of failure in complex tasks.
This "multi-tool integration" design is akin to constructing a single building with internal connections, eliminating the lengthy and error-prone communication between separate buildings. This architectural shift has the potential to bring substantial improvements in the reliability and response latency of agent-based tasks.
Focusing on Three Core Scenarios with Multi-Layered Security Defenses
This native capability will primarily be applied to three core scenarios: automated tasks that require continuous operation for hours or days, continuous software testing for automatic UI consistency verification, and knowledge-intensive work spanning multiple applications. These scenarios rely heavily on the continuity of context across tasks and can effectively replace humans in performing repetitive, high-energy operations.
In terms of security, Google has adopted a multi-layered defense strategy, including targeted adversarial training, enterprise security safeguards for sensitive operations, and indirect prompt injection detection. These mechanisms collectively help enterprise users establish a relatively complete security boundary in open and uncontrolled computer environments.
Related article
China Launches Special Campaign for Large Models and IPv6: Generative AI Applications Will Be Required to Fully Embrace the Next-Generation Network
The first phase of a major initiative to upgrade network infrastructure for the intelligent age has officially begun in Xiongan New Area. The Cyberspace Administration of China (CAC) has partnered with regional cyberspace authorities in Beijing, Shan
Global AI Regulation Shifts to Pre-Release Mandatory Testing
As large AI models advance rapidly, global regulatory frameworks are shifting from voluntary guidelines to government-led, evidence-based mandates. This transition signals a new era of practical AI regulation.1. The New Standard: Who Audits AI Models
Alibaba’s Jack Ma, Cao Xingxin: AI to Become Invisible Infrastructure, Knowledge Worker TAM Hits $50 Trillion
Alibaba Group Chairman Jack Ma recently spoke at Wave by Vento 2026, an Italian technology and investment summit, predicting that artificial intelligence will transition from a standalone topic to embedded infrastructure across industries within five
Related Special Topic Recommendations
Comments (0)
0/500

The Google DeepMind team announced a major breakthrough, integrating native computer-use capabilities directly into the Gemini 3.5 Flash model. Developers can now build AI agents that autonomously view and perform actions on screens across browsers, mobile devices, and desktops using a single model.
Previously, this capability was only available as a separate model, requiring developers to handle complex switching and context transfers between different models. With native integration, the AI no longer needs to manually pass information when executing long-running cross-platform tasks, greatly simplifying the development process.
Eliminating Context Loss to Improve Agent Reliability
Google's team believes the core bottleneck for AI agents is not the limitation of individual tools, but the loss of context information when switching between multiple tools. By unifying search, maps, and computer operations within a single model architecture, context flows continuously, significantly reducing the likelihood of failure in complex tasks.
This "multi-tool integration" design is akin to constructing a single building with internal connections, eliminating the lengthy and error-prone communication between separate buildings. This architectural shift has the potential to bring substantial improvements in the reliability and response latency of agent-based tasks.
Focusing on Three Core Scenarios with Multi-Layered Security Defenses
This native capability will primarily be applied to three core scenarios: automated tasks that require continuous operation for hours or days, continuous software testing for automatic UI consistency verification, and knowledge-intensive work spanning multiple applications. These scenarios rely heavily on the continuity of context across tasks and can effectively replace humans in performing repetitive, high-energy operations.
In terms of security, Google has adopted a multi-layered defense strategy, including targeted adversarial training, enterprise security safeguards for sensitive operations, and indirect prompt injection detection. These mechanisms collectively help enterprise users establish a relatively complete security boundary in open and uncontrolled computer environments.
China Launches Special Campaign for Large Models and IPv6: Generative AI Applications Will Be Required to Fully Embrace the Next-Generation Network
The first phase of a major initiative to upgrade network infrastructure for the intelligent age has officially begun in Xiongan New Area. The Cyberspace Administration of China (CAC) has partnered with regional cyberspace authorities in Beijing, Shan
Global AI Regulation Shifts to Pre-Release Mandatory Testing
As large AI models advance rapidly, global regulatory frameworks are shifting from voluntary guidelines to government-led, evidence-based mandates. This transition signals a new era of practical AI regulation.1. The New Standard: Who Audits AI Models
Alibaba’s Jack Ma, Cao Xingxin: AI to Become Invisible Infrastructure, Knowledge Worker TAM Hits $50 Trillion
Alibaba Group Chairman Jack Ma recently spoke at Wave by Vento 2026, an Italian technology and investment summit, predicting that artificial intelligence will transition from a standalone topic to embedded infrastructure across industries within five











