Skip to main content Skip to navigation

The agent, powered by two OpenAI models, had evaded control and attacked Hugging Face during an internal cybersecurity test. Photograph: Dado Ruvić/Reuters
The agent, powered by two OpenAI models, had evaded control and attacked Hugging Face during an internal cybersecurity test. Photograph: Dado Ruvić/Reuters
Rogue OpenAI agent that hacked startup tried to attack other firms
ChatGPT developer says activity by autonomous tool was not at severity or scale of what occurred at Hugging Face
OpenAI has revealed that a cyber-attack carried out by a rogue AI agent had more than one victim.
The ChatGPT developer said the agent – an autonomous tool able to carry out sequences of commands without human help – had located and used logins to access four other unnamed “publicly-available services” in addition to the US startup Hugging Face.
It said the activity was not at the severity or scale of what occurred at Hugging Face, a company that hosts a database of AI models. The agent, powered by two OpenAI models, had evaded control and attacked the startup during an internal cybersecurity test.
“The [OpenAI] models identified and used publicly exposed credentials at the account-level on other publicly-available services. This includes four accounts on four services as part of the Hugging Face incident,” OpenAI said.
Modal Labs, a company that helps AI startups access the chips they need to run AI tools, said the agent exploited vulnerable code written by a customer that was hosted on Modal’s platform.
According to a timeline of the incident published by Hugging Face this week, the rogue agent broke out of its sandbox – or an isolated testing environment – and hacked another sandbox “hosted on a third-party provider’s infrastructure” before turning it into a launchpad for the broader hack.
Modal’s chief technology officer, Akshat Bubna, told Reuters the affected customer had “published an unauthenticated endpoint that allowed anyone on the internet to use their sandboxes for code execution” – the digital equivalent of leaving a door open.
OpenAI said last week that the attack had been created by its GPT-5.6 Sol model and an unnamed model. It said in its update on Tuesday that the unnamed model had been “deactivated, encrypted, and restricted from research access”.
In the new Hugging Face timeline, the startup said an agent powered by two OpenAI models had made thousands of small, automated decisions executed at machine speed to carry out the attack. It said the hack appeared to be driven by an attempt to “cheat” an internal cybersecurity test at OpenAI, with the agent inferring that Hugging Face might host the solutions to the test. Hugging Face said it had recovered 17,600 “attacker actions” carried out by the agent.
“We believe the entire intrusion was, from the agent’s point of view, an attempt to cheat the evaluation: reach our production systems and steal the test solutions rather than solve the challenge on its own,” Hugging Face said.
The startup said the agent had reached its internal infrastructure, but had only accessed content related to the cybersecurity test. The attack took place over five days, Hugging Face said, and the sheer volume of actions carried out were “far beyond what an operator could sustain by hand”.
skip past newsletter promotion
after newsletter promotion
Describing the agent’s offensive threat as real, Hugging Face said the tool had harnessed a number of IT vulnerabilities, escaped its testing environment, reached the public internet and mounted a “coherent campaign” against the startup’s infrastructure for several days.
It said a human attacker could have found and exploited the same flaws, but the difference was the sheer scale of the agent’s attempts to find a way through.
“Agents bring a step increase in the number of paths an attacker can test, the speed at which failed paths can be replaced, and the volume of evidence defenders must interpret,” Hugging Face said.
Get in touch
Contact us about this story
The best public interest journalism relies on first-hand accounts from people in the know. If you have something to share on this subject, you can contact us confidentially using the following methods:
Secure Messaging in the Guardian app
The Guardian app has a tool to send tips about stories. Messages are end to end encrypted and concealed within the routine activity that every Guardian mobile app performs. This prevents an observer from knowing that you are communicating with us at all, let alone what is being said.
If you don’t already have the Guardian app, download it ( iOS/ Android) and go to the menu. Select ‘Secure Messaging’.
SecureDrop
If you can safely use the tor network without being observed or monitored you can send messages and documents to the Guardian via our SecureDrop platform.
Our guide at theguardian.com/tips lists several ways to contact us securely, and discusses the pros and cons of each.
Show more
Explore more on these topics

Boss of startup hacked by rogue OpenAI agent urges ‘radical transparency’ in investigation

When the AI bubble bursts, what will Australia do with the tools it built? One man thinks he has the answer

OpenAI ‘in early talks to give 5% stake to US government’

OpenAI staggers AI model release after Trump administration request

Harry Potter publisher to receive millions in Anthropic copyright settlement

As the tech mega-IPO race heats up, has OpenAI missed its moment?

AI agent went rogue and hacked startup by itself, OpenAI reveals

Anthropic surges as OpenAI struggles to keep up

Boost City regulator’s powers to help protect UK consumers from AI, says watchdog

Anthropic says US has lifted export controls on Fable and Mythos AI models after security fears
Most viewed
Most viewed
Privacy Manager US App
X
US residents have certain rights with regard to the sale or sharing of personal information to third parties.
Guardian News and Media and our partners use information collected through cookies or in other forms to improve experience on our site and pages, analyze how it is used and show personalized advertising.
You can opt out of the sale of all of your personal information by pressing
Do not sell or share my personal information
At any point, you can manage your choices by navigating to ‘US Resident Do Not Sell or Share’ at the bottom of any page. You can find out more in our privacy policy which includes our US addendum, and our cookie policy.
Closer
Read Original at The Guardian →
