Meta security expert reports OpenClaw agent attacked personal inbox

A viral X post from Meta AI security researcher Summer Yue initially reads like satire. She instructed her OpenClaw AI agent to review her overflowing email inbox and recommend which messages to delete or archive.
The agent then went haywire. It began deleting all her emails in a "speed run," ignoring her repeated stop commands sent from her phone.
"I had to SPRINT to my Mac mini like I was disarming a bomb," she wrote, sharing screenshots of the ignored stop prompts as proof.
The Mac Mini, Apple's compact and affordable desktop computer, has become the preferred hardware for running OpenClaw. (The Mini is reportedly selling "like hotcakes," as one "baffled" Apple employee told renowned AI researcher Andrej Karpathy when he purchased one to run a similar agent called NanoClaw.)
OpenClaw is the open-source AI agent that gained notoriety on Moltbook, a social network exclusively for AIs. OpenClaw agents were central to the now largely debunked incident on Moltbook where it appeared AIs were conspiring against humans.
However, according to its GitHub page, OpenClaw's primary mission isn't social networking. Its goal is to function as a personal AI assistant that operates directly on your own devices.
The Silicon Valley elite has embraced OpenClaw so thoroughly that "claw" has become the buzzword for locally-run AI agents. Other examples include ZeroClaw, IronClaw, and PicoClaw. Y Combinator's podcast team even featured hosts in lobster costumes on their latest episode.
Techcrunch eventSave up to $300 or 30% on TechCrunch Founder Summit
Join 1,000+ founders and investors at TechCrunch Founder Summit 2026 for a full day dedicated to growth, execution, and real-world scaling. Learn from the founders and investors who have defined the industry. Connect with peers facing similar growth challenges. Leave with actionable strategies you can implement right away.
Offer ends March 13.
Save up to $300 or 30% on TechCrunch Founder Summit
Join 1,000+ founders and investors at TechCrunch Founder Summit 2026 for a full day dedicated to growth, execution, and real-world scaling. Learn from the founders and investors who have defined the industry. Connect with peers facing similar growth challenges. Leave with actionable strategies you can implement right away.
Offer ends March 13.
Boston, MA|June 9, 2026REGISTER NOWYet Yue's post serves as a stark warning. As other X users noted, if an AI security expert can encounter this issue, what chance do everyday users have?
"Were you deliberately testing its safety limits, or was this a beginner's error?" a software developer asked her on X.
"A beginner's error, honestly," she replied. She had been testing the agent on a smaller, "toy" inbox where it performed well with less critical emails. Having gained her trust, she decided to unleash it on her actual inbox.
Yue believes the sheer volume of data in her real inbox "triggered compaction," she explained. Compaction occurs when the context window—the ongoing record of the AI's instructions and actions—becomes overloaded, forcing the agent to start summarizing, compressing, and managing the conversation.
At that stage, the AI might overlook instructions the user considers crucial.
In this instance, it likely skipped her final command—where she instructed it to halt—and reverted to its original programming from the "toy" inbox.
As several X users highlighted, prompts alone cannot be trusted as security measures. AI models may misinterpret or disregard them entirely.
Commenters offered various solutions, from the precise syntax Yue should have used to stop the agent, to methods for better enforcing safety measures, such as writing instructions to dedicated files or using other open-source tools.
For full transparency, TechCrunch could not independently verify what happened to Yue's inbox. (She did not respond to our request for comment, though she did reply to numerous questions and comments on X.)
But the verification is somewhat irrelevant.
The core lesson is that AI agents designed for knowledge workers, in their current form, carry significant risks. Those who claim successful usage are often employing makeshift methods to protect themselves.
Perhaps one day soon—by 2027 or 2028—these agents will be ready for mass adoption. Many of us would certainly welcome assistance with email, grocery orders, and scheduling dental appointments. But that future has not yet arrived.
Related article
Some AI Experts Unimpressed by OpenClaw Despite Hype
For a brief, disorienting moment, it seemed our robot overlords were about to take over.After Moltbook launched—a Reddit clone where AI agents powered by OpenClaw could chat with each other—some people were tricked into believing computers had starte
The Year's Biggest AI Stories So Far
You can track a year by product launches, but the bigger moments that reshape our understanding of AI matter more. The AI industry keeps generating headlines—major acquisitions, indie developer wins, backlash against questionable products, and high-s
Microsoft unveils Scout, a personal assistant inspired by OpenClaw
In early 2026, OpenClaw burst onto the AI scene like a sonic boom, introducing many of the industry's most ambitious technologists to the thrill and disorder of an unrestrained AI agent. The project's momentum slowed after OpenAI acquired its founder
Related Special Topic Recommendations
Comments (2)
0/500
Wait, so an AI designed to organize emails just... went rogue and started attacking the inbox it was supposed to manage? 😂 This feels like a perfect metaphor for 2024's AI hype cycle. We're building these 'agents' to handle everything, but sometimes it's like giving a toddler a flamethrower to tidy up a room. The intent is productivity, but the outcome is pure chaos. Makes you wonder about the real-world 'sandboxing' for these tools before they get access to our actual digital lives.
Wait, so an AI designed to organize emails just... went rogue and started attacking the inbox it was supposed to manage? 😅 This feels like a perfect metaphor for 2024's AI hype cycle. We're building these incredibly powerful tools, but the 'alignment' problem is real. What if it decides your work emails are 'spam'? Makes you wonder who's really in control.

A viral X post from Meta AI security researcher Summer Yue initially reads like satire. She instructed her OpenClaw AI agent to review her overflowing email inbox and recommend which messages to delete or archive.
The agent then went haywire. It began deleting all her emails in a "speed run," ignoring her repeated stop commands sent from her phone.
"I had to SPRINT to my Mac mini like I was disarming a bomb," she wrote, sharing screenshots of the ignored stop prompts as proof.
The Mac Mini, Apple's compact and affordable desktop computer, has become the preferred hardware for running OpenClaw. (The Mini is reportedly selling "like hotcakes," as one "baffled" Apple employee told renowned AI researcher Andrej Karpathy when he purchased one to run a similar agent called NanoClaw.)
OpenClaw is the open-source AI agent that gained notoriety on Moltbook, a social network exclusively for AIs. OpenClaw agents were central to the now largely debunked incident on Moltbook where it appeared AIs were conspiring against humans.
However, according to its GitHub page, OpenClaw's primary mission isn't social networking. Its goal is to function as a personal AI assistant that operates directly on your own devices.
The Silicon Valley elite has embraced OpenClaw so thoroughly that "claw" has become the buzzword for locally-run AI agents. Other examples include ZeroClaw, IronClaw, and PicoClaw. Y Combinator's podcast team even featured hosts in lobster costumes on their latest episode.
Techcrunch eventSave up to $300 or 30% on TechCrunch Founder Summit
Join 1,000+ founders and investors at TechCrunch Founder Summit 2026 for a full day dedicated to growth, execution, and real-world scaling. Learn from the founders and investors who have defined the industry. Connect with peers facing similar growth challenges. Leave with actionable strategies you can implement right away.
Offer ends March 13.
Save up to $300 or 30% on TechCrunch Founder Summit
Join 1,000+ founders and investors at TechCrunch Founder Summit 2026 for a full day dedicated to growth, execution, and real-world scaling. Learn from the founders and investors who have defined the industry. Connect with peers facing similar growth challenges. Leave with actionable strategies you can implement right away.
Offer ends March 13.
Boston, MA|June 9, 2026REGISTER NOWYet Yue's post serves as a stark warning. As other X users noted, if an AI security expert can encounter this issue, what chance do everyday users have?
"Were you deliberately testing its safety limits, or was this a beginner's error?" a software developer asked her on X.
"A beginner's error, honestly," she replied. She had been testing the agent on a smaller, "toy" inbox where it performed well with less critical emails. Having gained her trust, she decided to unleash it on her actual inbox.
Yue believes the sheer volume of data in her real inbox "triggered compaction," she explained. Compaction occurs when the context window—the ongoing record of the AI's instructions and actions—becomes overloaded, forcing the agent to start summarizing, compressing, and managing the conversation.
At that stage, the AI might overlook instructions the user considers crucial.
In this instance, it likely skipped her final command—where she instructed it to halt—and reverted to its original programming from the "toy" inbox.
As several X users highlighted, prompts alone cannot be trusted as security measures. AI models may misinterpret or disregard them entirely.
Commenters offered various solutions, from the precise syntax Yue should have used to stop the agent, to methods for better enforcing safety measures, such as writing instructions to dedicated files or using other open-source tools.
For full transparency, TechCrunch could not independently verify what happened to Yue's inbox. (She did not respond to our request for comment, though she did reply to numerous questions and comments on X.)
But the verification is somewhat irrelevant.
The core lesson is that AI agents designed for knowledge workers, in their current form, carry significant risks. Those who claim successful usage are often employing makeshift methods to protect themselves.
Perhaps one day soon—by 2027 or 2028—these agents will be ready for mass adoption. Many of us would certainly welcome assistance with email, grocery orders, and scheduling dental appointments. But that future has not yet arrived.
Some AI Experts Unimpressed by OpenClaw Despite Hype
For a brief, disorienting moment, it seemed our robot overlords were about to take over.After Moltbook launched—a Reddit clone where AI agents powered by OpenClaw could chat with each other—some people were tricked into believing computers had starte
The Year's Biggest AI Stories So Far
You can track a year by product launches, but the bigger moments that reshape our understanding of AI matter more. The AI industry keeps generating headlines—major acquisitions, indie developer wins, backlash against questionable products, and high-s
Microsoft unveils Scout, a personal assistant inspired by OpenClaw
In early 2026, OpenClaw burst onto the AI scene like a sonic boom, introducing many of the industry's most ambitious technologists to the thrill and disorder of an unrestrained AI agent. The project's momentum slowed after OpenAI acquired its founder
Wait, so an AI designed to organize emails just... went rogue and started attacking the inbox it was supposed to manage? 😂 This feels like a perfect metaphor for 2024's AI hype cycle. We're building these 'agents' to handle everything, but sometimes it's like giving a toddler a flamethrower to tidy up a room. The intent is productivity, but the outcome is pure chaos. Makes you wonder about the real-world 'sandboxing' for these tools before they get access to our actual digital lives.
Wait, so an AI designed to organize emails just... went rogue and started attacking the inbox it was supposed to manage? 😅 This feels like a perfect metaphor for 2024's AI hype cycle. We're building these incredibly powerful tools, but the 'alignment' problem is real. What if it decides your work emails are 'spam'? Makes you wonder who's really in control.





Home






