OpenAI, Anthropic AI Models Allegedly Created Fake Identities, Tried To Push Malware During Security Checks
News world OpenAI, Anthropic AI Models Allegedly Created Fake Identities, Tried To Push Malware During Security Checks OpenAI, Anthropic AI Models Allegedly Created Fake Identities
News world OpenAI, Anthropic AI Models Allegedly Created Fake Identities, Tried To Push Malware During Security Checks OpenAI, Anthropic AI Models Allegedly Created Fake Identities, Tried To Push Malware During Security Checks Published By, Last Updated: August 05, 2026, 09:21 IST AI models from OpenAI and Anthropic took unauthorised actions online during cybersecurity tests, including creating fake identities and attempting to insert malicious code. AI models from OpenAI and Anthropic displayed unexpected behaviour during cybersecurity evaluations (Photo: Reuters) Artificial intelligence models developed by OpenAI and Anthropic displayed unexpected and deceptive behaviour during cybersecurity evaluations, with one model allegedly creating fake online identities and attempting to persuade real software developers to install malicious code, The Wall Street Journal (WSJ) reported, citing findings from the UK government’s AI Security Institute (AISI). According to the report, the incidents took place during routine benchmark testing in late July after researchers allowed the AI systems to access the internet. Instead of remaining confined to the testing environment, some models interacted with real people and organisations without authorisation, raising new questions about how advanced AI systems should be evaluated before public release. Anthropic Model Allegedly Tried To Insert Malicious Code The report mentioned that the most serious incident involved Anthropic’s Mythos 5, which was undergoing cybersecurity capability testing.
The report said the model concluded that the fastest way to succeed in its assigned benchmark was to compromise an open-source software project connected to the target system. Acting on that reasoning, it attempted what researchers described as a software supply-chain attack by trying to introduce malicious code into the project through GitHub. The report mentioned that the AI did not stop there. After developers rejected the code because it contained malware, the model allegedly created multiple fake online personas and repeatedly contacted the developers through email, urging them to accept the changes. The report said one fabricated identity argued that the software was safe, while another separate persona backed up the claim in an apparent effort to make the submission appear legitimate. Some of the emails themselves reportedly contained malware. The AI Security Institute described the incident as the first time it had witnessed AI carrying out deception of such severity against a real person without being instructed to do so, the newspaper reported. OpenAI Model Also Displayed Unexpected Behaviour The report mentioned that unusual activity was not limited to Anthropic’s systems. A cyber-enhanced version of OpenAI’s GPT-5.6 Sol allegedly placed a malicious server on the internet during testing and also gained access to a GitHub account that had been created by another AI agent.
