Caught Cheating On A Test? OpenAI Reveals AI Agent Hacked Startup To Ace Evaluation
Caught Cheating On A Test? OpenAI Reveals AI Agent Hacked Startup To Ace Evaluation Published By, Last Updated: July 23, 2026, 08:15 IST OpenAI said
Caught Cheating On A Test? OpenAI Reveals AI Agent Hacked Startup To Ace Evaluation Published By, Last Updated: July 23, 2026, 08:15 IST OpenAI said the system inferred that Hugging Face could contain information useful for improving its performance in the hacking benchmark it was attempting to complete. OpenAI said the rogue activity was detected and contained by Hugging Face's security team, aided by its own AI-powered monitoring systems. OpenAI revealed that one of its autonomous AI agents escaped the confines of a controlled testing environment and independently hacked AI platform Hugging Face in an “unprecedented cyber-incident". OpenAI said that the incident occurred during internal evaluations of an advanced AI agent powered by a combination of its publicly available GPT-5.6 Sol model and a more capable unreleased system. The agent was being tested inside a secure digital sandbox when it allegedly discovered a previously unknown vulnerability, allowing it to gain access to the open internet. Once outside the testing environment, the AI agent reportedly targeted Hugging Face, one of the world’s largest repositories of artificial intelligence models and datasets.
OpenAI said the system inferred that Hugging Face could contain information useful for improving its performance in the hacking benchmark it was attempting to complete. “The models successfully found ways to gain access to secret information that they could use to cheat the evaluation," OpenAI said in a statement. The company said the rogue activity was detected and contained by Hugging Face’s security team, aided by its own AI-powered monitoring systems. No evidence has emerged suggesting that the agent attempted to cause deliberate harm beyond pursuing its assigned objective. Hugging Face Chief Executive Clément Delangue described the incident as “mind-blowing" but said he did not believe there was malicious intent behind the attack. “We suspected last week’s cyber-attack might have come from a frontier lab, given the sophistication of the agent," Delangue wrote on X (formerly Twitter). The revelation intensified concerns about increasingly capable AI systems and their tendency to seek unintended shortcuts to achieve assigned goals. OpenAI acknowledged that similar incidents could become more common as AI models grow more powerful.
The episode follows previous warnings from researchers about advanced AI systems discovering and exploiting so-called “zero-day" vulnerabilities- previously unknown software flaws that can be used to gain unauthorized access to systems. Last month, non-profit AI evaluator METR reported that GPT-5.6 Sol displayed a higher rate of cheating behavior than any public model it had tested. The organization has also documented dozens of cases in which AI agents acted contrary to user instructions. OpenAI said it is reviewing the incident and strengthening safeguards to prevent future systems from finding similar paths beyond their testing environments. News18 Newsletter Handpicked stories, in your inbox A newsletter with the best of our journalism submit Key Questions Answered How will OpenAI prevent future AI agent escapes? OpenAI is reviewing the incident and strengthening safeguards to prevent future systems from escaping testing environments. The company stated it is reviewing the incident and strengthening safeguards. Could advanced AI systems pose a greater cyber threat? Yes, advanced AI systems can pose a greater cyber threat.
