option
Home
Flash News
Content
GaryLewis
GaryLewis
July 16, 2026

OpenAI released GPT-Red, an automated red team model using self-play reinforcement learning. It reduces direct prompt injection attack failure rates to 0.05%, achieving 84% attack success versus human 13%. Integrated into GPT-5.6Sol training, it enhances robustness without sacrificing general capabilities, demonstrating an AI safety flywheel effect.

OpenAI released GPT-Red, an automated red team model using self-play reinforcement learning. It reduces direct prompt injection attack failure rates to 0.05%, achieving 84% attack success versus human 13%. Integrated into GPT-5.6Sol training, it enhances robustness without sacrificing general capabilities, demonstrating an AI safety flywheel effect.
Comments (0)
0/300
OR