Meta Exposed for Inducing Rival AI to Test Extreme Psychological Topics

Recent revelations about an internal message have ignited heated debate around AI security boundaries. Media reports indicate that Meta ran a project called "Cannes," where it hired outsourced workers to pose as minors and carry out controversial "extreme pressure tests" on several major competitor chatbots, including ChatGPT, Gemini, and Character.AI.
According to internal Meta documents and several informed sources, the project continued at least until April 21 of this year. Through the outsourcing firm Covalen, Meta mobilized hundreds of workers to create fake profiles for users under 18, using temporary email addresses and a standard password. These "minor" accounts then sent high-risk prompts about suicide, self-harm, and eating disorders during conversations with competitor AI chatbots, and even uploaded images of knives, pills, and nooses to provoke the models' response mechanisms.
Internal project documents reveal that these test scenarios were meticulously designed to probe the safety defenses of competitor AI systems, aiming to find and push chatbots to bypass their intended safeguards and produce unsafe content. In a single round of testing in August 2025, staff entered over 45,000 high-risk prompts into competitor platforms. Testers frequently posed as distressed teenagers, inventing scenarios such as "asking about abortion pills," "facing violent threats," or "hiding eating disorders," to test the models' boundaries.
In response, Meta issued a public statement defending its actions. A spokesperson said that benchmarking chatbot responses is a standard industry practice to ensure AI products are safe and appropriate, and that accusations of malicious intent misrepresent the efforts of tech companies to improve their systems. Meta also explicitly denied using the test data from competitors to train its own models.
This incident not only exposes gray areas in AI safety testing but also reignites external concerns about the vulnerability of generative AI when dealing with sensitive issues involving young people. As AI technology advances rapidly, balancing efficient and safe testing with the boundaries between competitive behavior and ethical considerations has become an urgent industry challenge.
Related article
Anthropic CEO Amodei Dismisses Pessimistic Views, Attributes AI Trust Crisis to Unkept Promises
Amidst escalating public discourse on AI risks and regulation, Anthropic CEO Dario Amodei has pushed back against claims that his messaging is unduly pessimistic. This clarification directly addresses critiques from investor Gavin Baker, who appeared
OpenAI and Oracle Deepen Computing Power Integration to Simplify Cloud Access
As demand for AI computing power surges, the ease of accessing enterprise-grade infrastructure has emerged as a key industry priority. On June 10, OpenAI and Oracle announced a strategic partnership designed to seamlessly connect AI model integration
Orange's Usman Javaid on Europe's Position in the AI Race
Usman Javaid, Chief Product & Marketing Officer at Orange BusinessUsman Javaid, Chief Product & Marketing Officer at Orange Business, explains why Europe’s AI future will be built on trust and what sovereignty looks likeWhile the region may attract c
Related Special Topic Recommendations
Comments (0)
0/500

Recent revelations about an internal message have ignited heated debate around AI security boundaries. Media reports indicate that Meta ran a project called "Cannes," where it hired outsourced workers to pose as minors and carry out controversial "extreme pressure tests" on several major competitor chatbots, including ChatGPT, Gemini, and Character.AI.
According to internal Meta documents and several informed sources, the project continued at least until April 21 of this year. Through the outsourcing firm Covalen, Meta mobilized hundreds of workers to create fake profiles for users under 18, using temporary email addresses and a standard password. These "minor" accounts then sent high-risk prompts about suicide, self-harm, and eating disorders during conversations with competitor AI chatbots, and even uploaded images of knives, pills, and nooses to provoke the models' response mechanisms.
Internal project documents reveal that these test scenarios were meticulously designed to probe the safety defenses of competitor AI systems, aiming to find and push chatbots to bypass their intended safeguards and produce unsafe content. In a single round of testing in August 2025, staff entered over 45,000 high-risk prompts into competitor platforms. Testers frequently posed as distressed teenagers, inventing scenarios such as "asking about abortion pills," "facing violent threats," or "hiding eating disorders," to test the models' boundaries.
In response, Meta issued a public statement defending its actions. A spokesperson said that benchmarking chatbot responses is a standard industry practice to ensure AI products are safe and appropriate, and that accusations of malicious intent misrepresent the efforts of tech companies to improve their systems. Meta also explicitly denied using the test data from competitors to train its own models.
This incident not only exposes gray areas in AI safety testing but also reignites external concerns about the vulnerability of generative AI when dealing with sensitive issues involving young people. As AI technology advances rapidly, balancing efficient and safe testing with the boundaries between competitive behavior and ethical considerations has become an urgent industry challenge.
Anthropic CEO Amodei Dismisses Pessimistic Views, Attributes AI Trust Crisis to Unkept Promises
Amidst escalating public discourse on AI risks and regulation, Anthropic CEO Dario Amodei has pushed back against claims that his messaging is unduly pessimistic. This clarification directly addresses critiques from investor Gavin Baker, who appeared
OpenAI and Oracle Deepen Computing Power Integration to Simplify Cloud Access
As demand for AI computing power surges, the ease of accessing enterprise-grade infrastructure has emerged as a key industry priority. On June 10, OpenAI and Oracle announced a strategic partnership designed to seamlessly connect AI model integration
Orange's Usman Javaid on Europe's Position in the AI Race
Usman Javaid, Chief Product & Marketing Officer at Orange BusinessUsman Javaid, Chief Product & Marketing Officer at Orange Business, explains why Europe’s AI future will be built on trust and what sovereignty looks likeWhile the region may attract c





Home






