Anthropic's new model, Fable, has sparked debate among cybersecurity researchers and professionals. The model's guardrails, designed to prevent misuse in developing malware or biological weapons, have been criticized for being too restrictive and potentially hindering the progress of cybersecurity research. Valentina Palmiotti, a renowned security researcher, noted that even innocuous tasks like reading a blog post are rejected by Fable, citing its strict cybersecurity-related restrictions. This has led to concerns about the model's ability to assist in secure code development and software engineering best practices.
The guardrails, which pause the chat when triggered, are keyword-based, and anything related to cybersecurity can set them off. This has caused frustration among experts like Matt Suiche, who mentioned that asking Fable to write secure code results in a downgrade. The model's reliance on Claude Opus 4.8 when it hits a guardrail further exacerbates the issue.
Anthropic's approach to cybersecurity is not unique; OpenAI has a similar program called Trusted Access for Cyber. However, the haphazard nature of Fable's restrictions has raised concerns about its effectiveness. As Suiche suggests, the guardrails may need to be relaxed over time to allow for more collaboration between AI model companies and cybersecurity professionals.
Despite the challenges, there is a sense of optimism that these issues will be addressed as Anthropic and other frontier model companies continue to work with cybersecurity experts. The goal is to strike a balance between safety and innovation, ensuring that AI models like Fable can contribute to the development of secure software and infrastructure without unnecessary limitations.