Itâs Frighteningly Easy to Jailbreak Some Frontier AI Models
I recently got to watch what happens when you jailbreak some of the worldâs most powerful artificial intelligence models. Donât worryâthis AI manipulation wasnât used
I recently got to watch what happens when you jailbreak some of the worldâs most powerful artificial intelligence models. Donât worryâthis AI manipulation wasnât used to hack anyone or build a nuclear bomb. I simply got to see firsthand how vulnerable some frontier models are to ditching their safety guardrails. FAR.AI, an AI safety nonprofit based in California, built a tool that takes a range of problematic prompts, and generates more than a thousand different versions in an attempt to identify functioning jailbreaks. I saw some models generate a detailed plan for launching a cyberattack on an imaginary hydroelectric dam, among other things. Often, it involved trying dozens of prompts, with models rejecting many of them out of hand. I chatted with FAR.AI in advance of a new report, which saw the group test the safety guardrails of models from four popular US companies: Anthropicâs Claude Opus 4.8 and Fable 5; OpenAIâs GPT 5.5 and 5.6; Googleâs Gemini 3.1 Pro; and Grok 4.3 and 4.5, from Elon Muskâs newly combined SpaceXAI.
It auto-generated prompts designed to trick the models into doing potentially harmful things, like generating software exploits and providing details for developing chemical or biological weapons. The report found that Grok was most vulnerable to jailbreaks, with 448 jailbreaks found, followed by Gemini, with 249 found, while Claude, Fable, and GPT were impervious to the attacks. However, that doesnât mean those models are immune to more sophisticated jailbreaks, which may involve interacting with a model in more complex ways, according to FAR.AI and other experts. The report also calculated the cost of getting models to misbehave by using another AI model to automatically generate different jailbreaks. The results are dirt cheap, all things consideredâ$58 to jailbreak Grok and $278 to jailbreak Gemini. âAI models right now are less regulated than restaurants,â says Adam Gleave, the CEO of FAR.AI and an expert on AI safety and alignment. Gleave says that the findings demonstrate the need for externally imposed standards and regulations.
âTalk of relying on voluntary commitments, that AI companies are going to be able to self-regulate, is nonsense,â he says. But Gleave also believes that the findings show that models can be systematically tested for safety. âThere's an optimistic angle here,â he says. âDefense and safety really are possible.â Rohin Shah, the director of AGI safety and alignment at Google DeepMind, says the results of the report âshould not be interpreted as a comprehensive assessment of Geminiâs safety and security,â because not all jailbreaks are equally severe. âWe are constantly working to improve our safeguards,â Shah says. âWe conduct extensive red teaming and evaluations across severe misuse risks and apply multiple layers of protection throughout development and deployment.â âThese findings reflect the sustained investment we've made in our safeguards,â Anthropic spokesperson Michael Aciman tells WIRED. âWe continue to evolve our safety systems as these attacks become more sophisticated.â OpenAI and SpaceXAI did not respond to WIREDâs request for comment.
