Cybersecurity Experts Criticize Guardrails on Anthropic’s Fable

Anthropic launched its newest model, Fable, on Tuesday, positioning it as a public, restricted iteration of its highly anticipated cybersecurity-focused model, Mythos.
However, the limitations have sparked dissatisfaction among several cybersecurity researchers and professionals who have voiced their concerns online.
“[Fable] blocks any request that might be remotely related to cybersecurity, including harmless tasks like reading a blog post,” stated Valentina “Chompie” Palmiotti, a prominent security researcher at IBM X-Force.
When a user prompt activates these safety filters, Fable halts the conversation, displaying a message that its “safety protocols flagged this input for cybersecurity or biological topics.”
These restrictions aim to mitigate the risk of Fable being utilized to create malware or exploit software—a persistent concern for Anthropic. The biological constraints stem from similar worries regarding the development of biological weapons.
Upon releasing Mythos in April, Anthropic limited access to a select group of companies and organizations under Project Glasswing, an initiative designed to deploy the model for securing critical software and infrastructure. Last week, Anthropic broadened Mythos access to hundreds of organizations across 15 countries.
Despite these good intentions, many cybersecurity experts remain frustrated by the seemingly arbitrary nature of the restrictions. Matt Suiche, a seasoned cybersecurity professional, told TechCrunch that “if you ask it to write secure code, it interprets this as cybersecurity-related work rather than software engineering best practices, resulting in a downgrade.” Fable is configured to revert to Claude Opus 4.8 when guardrails are triggered. “It appears to rely on keywords, so any term within the ‘cybersecurity’ lexicon activates the filters,” he noted.
Contact Us
Do you have additional insights on how hackers are leveraging AI, or how cybersecurity firms are integrating AI? We welcome your input. From a personal device and network, you can securely reach Lorenzo Franceschi-Bicchierai on Signal at +1 917 257 1382, via Telegram and Keybase @lorenzofb, or by email.“However, it is understandable given that we are still in the early stages and they are refining their guardrails. I am confident they will evolve as Anthropic and other frontier model developers collaborate more closely with the new generation of cybersecurity firms,” said Suiche, a technical staff member at Tolmo, an AI cybersecurity startup. “It is preferable to over-restrict initially and relax the guardrails over time.”
Another researcher complained on X that “even requesting a code review” activates Fable’s safety filters.
Anthropic did not immediately respond to a request for comment.
Beyond model-level guardrails, Anthropic requires cybersecurity professionals to apply for its Cyber Verification Program. Approved applicants face fewer restrictions when using Claude for cybersecurity tasks. OpenAI offers a comparable program known as Trusted Access for Cyber.
Related article
White House Ditches AI Label for ‘Super Intelligence’
Loading the player…This week, the White House gathered nearly every major tech CEO in one room — including Zuckerberg, Bezos, Musk, and Anthropic’s Dario Amodei — to sign an AI safety pledge that President Donald Trump described as “morally binding.”
Anthropic Unveils Sonnet 5.5 as Faster, Cheaper Work Partner
Amidst the ongoing competition among AI models, Anthropic has launched the latest iteration of Sonnet, its mid-range offering, promising substantially improved speed and lower costs compared to the previous version.Anthropic positions Sonnet 5.5 as a
Anthropic’s latest feud with the Trump admin may actually help it, sales data suggests
Anthropic is experiencing a remarkable month.According to Ramp, the AI company surpassed OpenAI in business spending market share for the first time, closing out May with a $65 billion raise at a $965 billion valuation. Shortly after, Anthropic filed
Related Special Topic Recommendations
Comments (0)
0/500

Anthropic launched its newest model, Fable, on Tuesday, positioning it as a public, restricted iteration of its highly anticipated cybersecurity-focused model, Mythos.
However, the limitations have sparked dissatisfaction among several cybersecurity researchers and professionals who have voiced their concerns online.
“[Fable] blocks any request that might be remotely related to cybersecurity, including harmless tasks like reading a blog post,” stated Valentina “Chompie” Palmiotti, a prominent security researcher at IBM X-Force.
When a user prompt activates these safety filters, Fable halts the conversation, displaying a message that its “safety protocols flagged this input for cybersecurity or biological topics.”
These restrictions aim to mitigate the risk of Fable being utilized to create malware or exploit software—a persistent concern for Anthropic. The biological constraints stem from similar worries regarding the development of biological weapons.
Upon releasing Mythos in April, Anthropic limited access to a select group of companies and organizations under Project Glasswing, an initiative designed to deploy the model for securing critical software and infrastructure. Last week, Anthropic broadened Mythos access to hundreds of organizations across 15 countries.
Despite these good intentions, many cybersecurity experts remain frustrated by the seemingly arbitrary nature of the restrictions. Matt Suiche, a seasoned cybersecurity professional, told TechCrunch that “if you ask it to write secure code, it interprets this as cybersecurity-related work rather than software engineering best practices, resulting in a downgrade.” Fable is configured to revert to Claude Opus 4.8 when guardrails are triggered. “It appears to rely on keywords, so any term within the ‘cybersecurity’ lexicon activates the filters,” he noted.
Contact Us
Do you have additional insights on how hackers are leveraging AI, or how cybersecurity firms are integrating AI? We welcome your input. From a personal device and network, you can securely reach Lorenzo Franceschi-Bicchierai on Signal at +1 917 257 1382, via Telegram and Keybase @lorenzofb, or by email.“However, it is understandable given that we are still in the early stages and they are refining their guardrails. I am confident they will evolve as Anthropic and other frontier model developers collaborate more closely with the new generation of cybersecurity firms,” said Suiche, a technical staff member at Tolmo, an AI cybersecurity startup. “It is preferable to over-restrict initially and relax the guardrails over time.”
Another researcher complained on X that “even requesting a code review” activates Fable’s safety filters.
Anthropic did not immediately respond to a request for comment.
Beyond model-level guardrails, Anthropic requires cybersecurity professionals to apply for its Cyber Verification Program. Approved applicants face fewer restrictions when using Claude for cybersecurity tasks. OpenAI offers a comparable program known as Trusted Access for Cyber.
White House Ditches AI Label for ‘Super Intelligence’
Loading the player…This week, the White House gathered nearly every major tech CEO in one room — including Zuckerberg, Bezos, Musk, and Anthropic’s Dario Amodei — to sign an AI safety pledge that President Donald Trump described as “morally binding.”
Anthropic Unveils Sonnet 5.5 as Faster, Cheaper Work Partner
Amidst the ongoing competition among AI models, Anthropic has launched the latest iteration of Sonnet, its mid-range offering, promising substantially improved speed and lower costs compared to the previous version.Anthropic positions Sonnet 5.5 as a
Anthropic’s latest feud with the Trump admin may actually help it, sales data suggests
Anthropic is experiencing a remarkable month.According to Ramp, the AI company surpassed OpenAI in business spending market share for the first time, closing out May with a $65 billion raise at a $965 billion valuation. Shortly after, Anthropic filed





Home






