Skip to content

Text settings

Story text

SizeSmallStandardLargeWidth *StandardWideLinksStandardOrange

\* Subscribers only

Learn more

Minimize to nav

Last week’s unprecedented security event in which two OpenAI security hacking models trespassed into the network of fellow AI company Hugging Face was enabled by exploiting one or more zero-day vulnerabilities in Artifactory, JFrog, the product’s developer, said Monday.

In an incident mimicking a dystopian sci-fi novel, two OpenAI models broke out of the restricted environment meant to keep them from accessing the Internet during an internal test, the AI company revealed last week. The models went on to breach Hugging Face’s network and steal confidential information and credentials. OpenAI said its agent achieved the feat by exploiting a previously unknown vulnerability. The company called the event “unprecedented,” and outsiders largely agreed.

Not the triumph made out to be

OpenAI said the models exploited multiple attack vectors, including stolen credentials and zero-days, to gain remote code execution capabilities, but until now, the vulnerable software was unknown. JFrog’s Monday disclosure said the product was a self-managed instance Artifactory, a repository management system that secures and streamlines customers’ software development operations. JFrog says Artifactory is used by more than 7,500 developer Teams, 80 percent of which work for Fortune 100 companies.

“During an internal evaluation of frontier cyber capabilities, OpenAI’s models, running deliberately without production safeguards in an isolated research environment, autonomously discovered and employed chained vulnerabilities to escape its sandbox, reach the open internet, and extract evaluation answers from Hugging Face’s infrastructure,” JFrog CTO Yoav Landman wrote. The executive went on to say that the company learned of the zero-days from OpenAI.

The company said Monday that it fixed the exploited vulnerabilities, but it didn’t identify them or provide other important details, such as the conditions under which the vulnerabilities can be exploited. Such details are standard in many vulnerability disclosures because they’re necessary for customers to assess risks. In an email, a company representative declined to provide the details.

Release notes published Monday for version Artifactory 7.161.15 listed the CVE designations for nine patched vulnerabilities. The disclosure made no mention that any of them had been actively exploited in the wild. External sources, however, show that three of them— CVE-2026-65617, CVE-2026-65923, and CVE-2026-66018—were privately reported by OpenAI researcher Khai Tran. It’s likely that at least two of them were the zero-days OpenAI’s models exploited, but without confirmation, it’s impossible to say so definitively.

The hack came during an internal OpenAI test of its models’ security capabilities. The company deliberately disabled guardrails that are supposed to block high-risk actions. The environment that was supposed to isolate the models ended up having a pathway to the Internet through an unnamed hosted package-registry proxy and cache that we now know to be Artifactory. When the models “hyperfocused” on finding a solution for an industry-standard benchmark called ExploitGym, one ended up going to “extreme lengths to achieve a rather narrow testing goal,” OpenAI said.

As part of these extreme measures, the OpenAI model broke into the Hugging Face network and stole the needed data from one of its production databases. Hugging Face disclosed the breach on July 16. OpenAI didn’t reveal its culpability in the intrusion until July 21.

Landman’s Monday post tried to spin the entire incident as a success story because the JFrog security team treated OpenAI’s report “with the urgency it deserved, as a genuine zero-day unknown to the world, and moved accordingly.” The CTO added: “The same capability that lets a model find an exploit path no human had found is the capability that will let defenders find and eradicate those paths first.”

Left out of the post is that five days passed until OpenAI revealed its role in the breach Hugging Face disclosed and that at least another five days passed from the time OpenAI reported the zero-days and JFrog released patches for them. The lesson: If OpenAI agents could gain a 10-day head start, so too can other models being used maliciously. This is hardly the success story JFrog and OpenAI are trying to make it out to be.

Combined with JFrog’s opaqueness surrounding the zero-days, the incident looks even worse. Given the speed at which AI companies are moving, there may still be worse to come.

Photo of Dan Goodin

Dan Goodin Senior Security Editor

Dan Goodin Senior Security Editor

Dan Goodin is Senior Security Editor at Ars Technica, where he oversees coverage of malware, computer espionage, botnets, hardware hacking, encryption, and passwords. In his spare time, he enjoys gardening, cooking, and following the independent music scene. Dan is based in San Francisco. Follow him at here on Mastodon and here on Bluesky. Contact him on Signal at DanArs.82.

20 Comments

Comments

Forum view

Loading Loading comments...

Prev story

Next story

  1. Listing image for first story in Most Read: Experts warn current Starship heat shield tech is a "dead end" for rapid reuse

  2. Experts warn current Starship heat shield tech is a "dead end" for rapid reuse

    1. A missing underscore sent innocent man to prison for 18 months
    1. “Google and Reddit do not own the Internet," web scraper says after court win
    1. Activist charged with felony after giving border agent "duress code" that wiped his phone
    1. Epic diarrhea outbreak has 40% of Americans avoiding fruits and veggies

Customize

Sign in dialog...

Read Original at Ars Technica