option
Home
Flash News
Content
BillyAnderson
BillyAnderson
June 4, 2026

Safety researcher Kasra Rahjerdi tested LLMs on exploiting a vulnerable app with exposed Google backend credentials. GPT-5.5 solved 7/10 runs, highest success rate, but cost $9.46 per success. DeepSeek V4 Pro succeeded 3/10 at $0.62 per success, one-fifteenth the cost. Gemini failed due to rejection mechanisms.

Safety researcher Kasra Rahjerdi tested LLMs on exploiting a vulnerable app with exposed Google backend credentials. GPT-5.5 solved 7/10 runs, highest success rate, but cost $9.46 per success. DeepSeek V4 Pro succeeded 3/10 at $0.62 per success, one-fifteenth the cost. Gemini failed due to rejection mechanisms.
Comments (0)
0/300
OR