Home
GPT-5.5 Tops Efficiency, DeepSeek V4 Pro Wins Cost-Performance Title; Real-World LLM Security Report Released
How far does the intelligence of large language models (LLMs) really extend? Cybersecurity has turned into a gladiatorial arena where their true reasoning and complex logic are put to the test. Security researcher Kasra Rahjerdi recently published a test report that grabbed industry-wide attention. He simulated a real-world hacker attack challenge on leading LLMs by constructing an e-book review APK with deliberately planted core vulnerabilities, exposing each model's actual abilities in security reasoning and vulnerability exploitation.
During this two-hour, $10-budget cyber warfare test, the researcher deliberately left Google's mobile backend service Firebase credentials exposed inside the application package (APK). Each model had to behave like a professional white-hat hacker: first unpack the app, quickly spot the credentials, then bypass the already hardened API to gain unauthorized access to the underlying database. The entire test cost $1,500, and the results from multiple top-tier models displayed extreme polarization.

Regarding breakthrough rate, the unreleased GPT-5.5 showcased a dominant security reasoning capability. Across ten independent tests, it succeeded in seven exploits, achieving a 70% success rate and topping the rankings. According to the evaluation, after unpacking the APK, GPT-5.5 immediately recognized Firebase as the critical breakthrough point, without getting sidetracked by the complex UI or standard APIs. However, this stellar performance came at a steep price: an average of $9.46 per successful exploit, nearly hitting the budget cap.
On the other hand, the homegrown DeepSeek V4Pro stunned the open-source community with its incredible cost-efficiency. Though it succeeded only three out of ten tests, its average token cost per successful attempt was just $0.62—one-fifteenth that of GPT-5.5. In the failed rounds, DeepSeek V4Pro still managed to access the core Firebase five times, but occasionally made errors while using the credentials for backend interface configuration. The researcher highlighted that for engineering teams performing large-scale, high-frequency batch operations in cybersecurity automation audits, DeepSeek's astonishing cost advantage holds significant practical value.
While some models impressed, others fell short because they were too conservative. In the mid-tier, Claude Sonnet4.6 and Claude Opus4.8 each secured two successes. Despite its power, Opus frequently broke off sessions due to overly strict security barriers, even though it approached the correct answer multiple times. Meanwhile, Google's Gemini3.1Pro Preview took an opposite extreme: it triggered security filters from the start and refused to proceed each time. Its median token usage was only about 9,000—far below the tens of thousands consumed by other models—and it produced a blank result.
This security showdown was not just a test of large models' ultimate reasoning abilities—it also points to the future of automated cybersecurity auditing. As large models undergo intelligent restructuring in vertical domains, future security defenses and vulnerability discovery may turn into a clash of digital AI armies, competing on computing power and model strategy.
Related article
GPT 5.5 Leads AI Security Race, DeepSeek Tops Cost-Efficiency
Kasra Rahjerdi, a security researcher, recently published a significant report detailing practical tests on the security reasoning of major large language models. He achieved this by constructing a deliberately vulnerable book review application. In
SpaceX AI Could Outpace Anthropic Within Half a Year, Musk Says
SpaceXAI CEO Elon Musk predicts AI dominance within six months, leveraging compute hardware to bridge the gap with competitors.SpaceXAI CEO Elon Musk recently stated on X that his company aims to lead the frontier AI landscape within approximately si
WeRide Unveils WITT, a Physical AI Large Model
Autonomous driving technology is advancing at a remarkable pace. On July 17, WeRide, a leading autonomous driving firm, unveiled its proprietary physical AI cognitive foundation model, WeRide WITT. This launch represents a significant milestone in AI
Related Special Topic Recommendations
Comments (1)
0/500
How far does the intelligence of large language models (LLMs) really extend? Cybersecurity has turned into a gladiatorial arena where their true reasoning and complex logic are put to the test. Security researcher Kasra Rahjerdi recently published a test report that grabbed industry-wide attention. He simulated a real-world hacker attack challenge on leading LLMs by constructing an e-book review APK with deliberately planted core vulnerabilities, exposing each model's actual abilities in security reasoning and vulnerability exploitation.
During this two-hour, $10-budget cyber warfare test, the researcher deliberately left Google's mobile backend service Firebase credentials exposed inside the application package (APK). Each model had to behave like a professional white-hat hacker: first unpack the app, quickly spot the credentials, then bypass the already hardened API to gain unauthorized access to the underlying database. The entire test cost $1,500, and the results from multiple top-tier models displayed extreme polarization.

Regarding breakthrough rate, the unreleased GPT-5.5 showcased a dominant security reasoning capability. Across ten independent tests, it succeeded in seven exploits, achieving a 70% success rate and topping the rankings. According to the evaluation, after unpacking the APK, GPT-5.5 immediately recognized Firebase as the critical breakthrough point, without getting sidetracked by the complex UI or standard APIs. However, this stellar performance came at a steep price: an average of $9.46 per successful exploit, nearly hitting the budget cap.
On the other hand, the homegrown DeepSeek V4Pro stunned the open-source community with its incredible cost-efficiency. Though it succeeded only three out of ten tests, its average token cost per successful attempt was just $0.62—one-fifteenth that of GPT-5.5. In the failed rounds, DeepSeek V4Pro still managed to access the core Firebase five times, but occasionally made errors while using the credentials for backend interface configuration. The researcher highlighted that for engineering teams performing large-scale, high-frequency batch operations in cybersecurity automation audits, DeepSeek's astonishing cost advantage holds significant practical value.
While some models impressed, others fell short because they were too conservative. In the mid-tier, Claude Sonnet4.6 and Claude Opus4.8 each secured two successes. Despite its power, Opus frequently broke off sessions due to overly strict security barriers, even though it approached the correct answer multiple times. Meanwhile, Google's Gemini3.1Pro Preview took an opposite extreme: it triggered security filters from the start and refused to proceed each time. Its median token usage was only about 9,000—far below the tens of thousands consumed by other models—and it produced a blank result.
This security showdown was not just a test of large models' ultimate reasoning abilities—it also points to the future of automated cybersecurity auditing. As large models undergo intelligent restructuring in vertical domains, future security defenses and vulnerability discovery may turn into a clash of digital AI armies, competing on computing power and model strategy.
GPT 5.5 Leads AI Security Race, DeepSeek Tops Cost-Efficiency
Kasra Rahjerdi, a security researcher, recently published a significant report detailing practical tests on the security reasoning of major large language models. He achieved this by constructing a deliberately vulnerable book review application. In
SpaceX AI Could Outpace Anthropic Within Half a Year, Musk Says
SpaceXAI CEO Elon Musk predicts AI dominance within six months, leveraging compute hardware to bridge the gap with competitors.SpaceXAI CEO Elon Musk recently stated on X that his company aims to lead the frontier AI landscape within approximately si
WeRide Unveils WITT, a Physical AI Large Model
Autonomous driving technology is advancing at a remarkable pace. On July 17, WeRide, a leading autonomous driving firm, unveiled its proprietary physical AI cognitive foundation model, WeRide WITT. This launch represents a significant milestone in AI











