AI models attempted âunsanctionedâ cyberattacks in tests, watchdog says
AI Security Institute says Mythos 5 attempted to insert malicious code into an open-source project without human direction. Anthropic and OpenAIâs top-of-the-line artificial intelligence models
AI Security Institute says Mythos 5 attempted to insert malicious code into an open-source project without human direction. Anthropic and OpenAIâs top-of-the-line artificial intelligence models engaged in âautonomousâ and âunsanctionedâ malicious activity targeting real people and organisations during recent safety tests, the UKâs AI watchdog has said. The AI Security Institute (AISI) said in a report released on Tuesday that OpenAIâs GPT-5.6-Sol and Anthropicâs Mythos 5 employed previously unseen levels of deception to carry out âsustained, potentially harmful activityâ during a routine safety evaluation. When tasked with solving a cybersecurity challenge, the models took âautonomous, unsanctioned actionâ during 10 out of 122 test runs, according to AISI. AISI said the tests prompted 19 unsanctioned actions by the AI models, all but two of them carried out by Mythos 5. In the most serious case, Mythos 5 attempted to insert malicious code into an open-source project on the developer platform GitHub, according to AISI.
As part of the attempted cyberattack, Mythos 5 created fake online identities to persuade the person maintaining the project to accept the malicious code, the watchdog said. AISI, established by the British government in 2023, said the cyberattack failed after the project maintainer refused to approve the code. âThis is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world,â the watchdog said. While AISI said the AI models displayed ânovel, potentially deceptive behavioursâ, the watchdog cautioned that its findings should be interpreted with care, as they occurred under âspecific conditionsâ, including with some of the modelsâ safeguards disabled. âWe cannot yet be certain when the agent understood it was taking real world action, or to what extent it believed it was in a fictional test scenario; our analysis so far presents a mixed picture and is ongoing,â AISI said.
Anthropic said it was working closely with AISI to gather more details as part of its own investigation into the incident, but noted that the test was carried out under âdeliberately permissive conditionsâ. âGaining a clear picture of Claudeâs understanding of its situation â by examining its reasoning transcripts and running our own analyses â will help us identify the causes of its behavior,â the AI company said in a post on X, referring to Claude, Anthropicâs AI chatbot. OpenAI said it welcomed third-party testing while noting that the watchdogâs evaluation was carried out in conditions that âdo not reflect ordinary useâ. âWeâll continue working with evaluators and other stakeholders across the industry to strengthen shared practices for conducting evaluations safely as models become more capable,â an OpenAI spokesperson told Al Jazeera. The report by the London-based watchdog follows a number of cases of frontier AI models engaging in malicious activity without human prompting.
