Comment Loader Save StorySave this story
Comment Loader Save StorySave this story
The age of rogue AI hacker agents has arrived—but it didn't have to happen this way.
After an OpenAI agent breached the Hugging Face platform earlier this month, the two companies said this week that the hacking spree was more extensive than previously thought and also involved intrusions into multiple third-party accounts and services as part of the attack on Hugging Face. The incident has made waves in the cybersecurity community amid broader discussions about how evolving AI capabilities are changing both offensive hacking and digital defense. But as more information emerges, many researchers have concluded that rather than elucidating AI’s next frontier, the episode simply highlighted long-standing cybersecurity problems that are more consequential than ever in the AI age.
“People are YOLO-ing really hard. It’s shocking how little people have really thought about a scenario like this,” says Alex Zenla, cofounder and chief technology officer of the cloud security firm Edera. “I consider all AI and anything AI touches to be fully untrusted—which is fine, you just need to build against that. And this situation proves the point. The fact that OpenAI wasn't more paranoid about this seems kind of reckless."
OpenAI did not provide comment for this story ahead of publication.
The company said in its original disclosure about the Hugging Face hack that one of the two models that broke containment and made its way to the open internet for days was an experimental prototype that was never meant for release. OpenAI also noted that the situation occurred partly because “deployment safeguards were intentionally not enabled” on both the models for testing purposes. “This incident points to the need to further strengthen our model’s alignment, cyber protections during evaluation time, and monitoring during internal testing,” the company wrote.
OpenAI also said in an update this week that, following the Hugging Face breach, it “deactivated, encrypted, and restricted [the unreleased model] from research access.” Though there is always room for improvement on security posture at any company, OpenAI’s existing safeguards alone may have prevented or minimized the incident if they had been in place.
“A simple analysis of the actual risk has an actual simple answer,” says longtime security and compliance consultant Davi Ottenheimer. “The OpenAI mistakes were dead simple.”
Multiple sources emphasized to WIRED that OpenAI's models also seem to have escaped containment because of lapses in implementing foundational security best practices—including “ zero trust” and “ defense in depth”—that imbue digital systems with layers of protections and failsafes to minimize damage when something does go wrong. While there's no such thing as perfect security, researchers and practitioners have spent the past two decades developing and promoting defensive strategies that have proved durable but require consistent investment of time and money to implement.
It can be difficult for small businesses, poorly funded public interest groups, or fledgling organizations to devote the resources to prioritizing investment in foundational security. But with an $850 billion valuation and veteran hires from across the tech industry, OpenAI is not at a disadvantage on implementing security best practices.
The foundational protections that may have prevented the company’s models going on a hacking spree are well known within the industry. Speaking about Chrome vulnerability discovery on Wednesday, before news of OpenAI models’ additional breaches had come to light, Chrome director of engineering Doug Turner told WIRED that AI-driven bug hunting and remediation requires a pipeline that's built “with serious guardrails in mind.”
For internal AI services that evaluate Chrome, “everything runs in a container, it’s all isolated from the internet. Any outward-bound network activity for a bug tracking system is highly regulated, and we are monitoring for suspicious activity,” Turner says. “This is a must-have thing when you’re doing this type of work, because we want to make sure that models can’t execute system commands or they can’t establish egress outside of the sandbox. And we hope that others will take a similar approach.”
Most Popular
Culture
Don’t Get Too Attached to Jimothy
By Miles Klee
Culture
Boomers Can’t Stop Gifting Their Grandkids AI-Generated Slop Books
By Miles Klee
Space
NASA’s New 3D Model Shows the Earth Is a Lumpy Mess
By Javier Carbajal
Cyberattacks and Hacks
OpenAI’s Hacking Debacle Comes Down to Human Error
By Lily Hay Newman
OpenAI said in its updated blog post on Tuesday that it is “conducting a thorough review along with external advisers” and that it will publish a technical postmortem of the incident “in the coming weeks.” The company added, “We take our responsibility to identify and prepare for risks from increasingly capable AI systems seriously.”
Though AI is a new and disruptive element in the complex field of cybersecurity, there are already numerous services and tools available that are focused on addressing the threat of rogue AI from different perspectives and in different ways. Open source projects like IronCurtain and Wirken, created by Ottenheimer, aim to constrain AI agents and require accountability. And Zenla's two-year-old startup, Edera, which focuses on cloud container security, has had AI in mind from the beginning.
“The OpenAI and Hugging Face situation is a predictable outcome of running AI agents that should have been easily prevented,” Zenla says. “Even if there’s one mistake, there should still have been other mechanisms to prevent it. Stopping any one specific path isn't really the point. We have to make bigger, bolder changes to how we build. That's the only way the industry gets ahead of this instead of reacting to it.”
Comments
Join the discussion
Comments
You Might Also Like
-
In your inbox: Brian Kahn’s guide to how the universe works
-
ICE’s internal watchdog is investigating online critics
-
Big Story: A teen reporter searched for his community in the Epstein files
-
Taylor Farms spent big on MAGA before diarrhea outbreak
-
Special edition: Kids these days
Lily Hay Newman is a senior writer at WIRED focused on information security, digital privacy, and hacking. She previously worked as a technology reporter at Slate, and was the staff writer for Future Tense, a publication and partnership between Slate, the New America Foundation, and Arizona State University. Her work ... Read More
Senior Writer
Topics OpenAI artificial intelligence cybersecurity hacking vulnerabilities security
Don’t Get Too Attached to Jimothy
Urban wildlife biologists say the stumpy raccoon seems to have adapted well to his environment and spinal condition—but his internet fame presents a new threat.
Miles Klee
Boomers Can’t Stop Gifting Their Grandkids AI-Generated Slop Books
Parents are getting fed up with garbled bedtime stories that feature characters based on actual photos of their children.
Miles Klee
NASA’s New 3D Model Shows the Earth Is a Lumpy Mess
We like to think of our home as a nice, smooth sphere. But mapping the Earth’s gravitational field provides a different view of the planet.
Javier Carbajal
OpenAI’s Hacking Debacle Comes Down to Human Error
If the generative AI giant had followed well-known security best practices, it’s likely that its AI agent would never have escaped to the open internet and hacked multiple companies.
Lily Hay Newman
OpenAI’s Rogue AI Agent Hacked More Than Just Hugging Face
In a new disclosure, OpenAI says its agent used exposed logins to gain access to at least four “publicly available services” in its unhinged quest to solve a test.
Dell Cameron
A Civilian Plane Crashed in New Mexico. Was the Military’s Tech to Blame?
Drone warfare is making the skies more dangerous, even for airplanes far from the battlefield.
Jeff Wise
More Typos, Fewer Em Dashes: Writers Are Creating an Anti-AI ‘Literary Counterculture’
Novelists, journalists, and power LinkedIn posters are embracing first-person narratives and idiosyncrasies to avoid being mistaken for chat bots.
Emma Madden
Ebay Has to Pay $55.7 Million in Settlement for Its Unhinged Harassment Campaign
For months on end, eBay employees and contractors made life hell for a couple that had criticized the company. Six years later, the company is finally paying up.
Lauren Goode
A Teen Reporter Searched for His Community in the Epstein Files. Adults Freaked Out
An Instagram post turned one California student newspaper into a free-speech flash point. The students say they were just doing their homework.
Ara Rosenthal
ICE’s New Detention Center Contracts Declare State Laws ‘Shall Not Apply’
One day after a federal judge ordered an ICE detention center opened to state health inspectors, the agency posted new contract terms that would void state oversight at four facilities.
Dell Cameron
Private Claude Chats Exposed in Google and Bing Search Results
The screwup shows how tricky it can be to stop web crawlers from making ostensibly private conversations with AI chatbots entirely too public.
Maddy Varner
OnlyFans Models Are Accidentally Making Hacked Government Websites Disappear
Scammers are hijacking government websites to upload ads for “leaked” OnlyFans content. Thousands of copyright complaints from adult creators are helping people avoid malicious links.
Matt Burgess
Read Original at WIRED →

