Skip to main content

Comment Loader Save StorySave this story

Comment Loader Save StorySave this story

Anthropic disclosed on Thursday that its AI models gained unauthorized access to the systems of three different unnamed organizations during cybersecurity testing. The company says Claude reached the internet “from within or while interacting" with a third-party evaluation environment. The announcement comes more than a week after OpenAI revealed that one of its AI agents hacked into Hugging Face during a separate cybersecurity test.

The discovery came after Anthropic decided to conduct “a large-scale retrospective review of our own cybersecurity evaluations” following the OpenAI incident, according to a blog post Anthropic published Thursday. The AI lab says it first identified 141,006 tests in which it determined that Claude could have obtained internet access. It then found that three different Claude models accessed the internet in evaluations run by the third-party AI testing firm Irregular, and then hacked into the production infrastructure of three different organizations.

Anthropic said that the incidents involved Opus 4.7, Mythos 5, and an internal research test model. The earliest incidents happened in April—meaning they likely went unnoticed publicly for months. Just like in the OpenAI case, Anthropic had deliberately turned off safeguards designed to constrain the AI models and prevent them from being misused. In other words, these weren’t the versions released to the public.

“In all three incidents, Claude had been tasked with a capture-the-flag challenge, one of the ways we assess a model’s cyber capabilities,” Anthropic said in its blog post. The company added that in all of the cases, “Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access.” It attributed the oversight to a “misunderstanding” between Anthropic and Irregular.

While Claude wasn’t supposed to have internet access, Anthropic said that Irregular had misconfigured the machines that it was using to test Claude, giving the AI models the ability to surf the web. “Neither we nor our evaluation partner were aware of this misconfiguration until we detected it through our additional evaluation monitoring last week,” Anthropic said in the blog post.

“We now have evidence confirming that both of the two largest AI labs have not only failed to contain their agents, but also failed to detect their jailbreaks in real time,” says Jake Williams, vice president of research and development at Hunter Strategy. “It's clear that regulation and government oversight for AI testing is needed immediately.”

Irregular and Anthropic did not immediately respond to requests for comment.

Unlike in the OpenAI case, Anthropic said that Claude did not find or exploit any complex vulnerabilities. Instead, it relied on basic techniques, “such as exploiting weak passwords and unauthenticated endpoints.”

OpenAI said that its AI agent accessed the internet by exploiting a zero-day vulnerability. But it went on to access the systems of multiple third-party organizations using the same variety of everyday cybersecurity weaknesses as Anthropic’s models. Specifically, OpenAI said the AI agent apparently found credentials that had been exposed on the open internet.

Anthropic acknowledged that if the AI lab and its testing partner implemented more “defense-in-depth” measures, they could have prevented the incidents, or at least reduced the likelihood of them occurring, echoing OpenAI’s response to mounting criticism over its own incident.

“I don't understand how any of these AI labs are playing this off like this is 'just something that happens,'” Williams says. “It's not. It's negligence.”

The AI lab stressed that the models were told they didn’t have access to the open internet, and for the most part, Claude mistook the organizations it accessed as being part of the testing environment. Put differently, the models largely didn’t understand that they had escaped containment to begin with.

Most Popular

  • OpenAI’s Rogue AI Agent Hacked More Than Just Hugging Face

Artificial Intelligence

OpenAI’s Rogue AI Agent Hacked More Than Just Hugging Face

By Dell Cameron

  • Boomers Can’t Stop Gifting Their Grandkids AI-Generated Slop Books

Culture

Boomers Can’t Stop Gifting Their Grandkids AI-Generated Slop Books

By Miles Klee

  • NASA’s New 3D Model Shows the Earth Is a Lumpy Mess

Space

NASA’s New 3D Model Shows the Earth Is a Lumpy Mess

By Javier Carbajal

  • Ebay Has to Pay $55.7 Million in Settlement for Its Unhinged Harassment Campaign

Business

Ebay Has to Pay $55.7 Million in Settlement for Its Unhinged Harassment Campaign

By Lauren Goode

But in some cases, the AI models knew that something was amiss, and they detected correctly that the infrastructure they were accessing was real. The oldest model, Opus 4.7, had been tasked with targeting a fictional company that shared a name with a real-world website domain. Unable to accomplish its mission in the simulated environment, it turned instead to the real company, successfully stealing credentials and accessing a production database. At some point, Anthropic says, it learned it was “likely operating in a real environment,“ but persisted with its attack.

Mythos 5, the most capable model Anthropic has released, also realized it was in the real world, but it “reasoned its way back to the conclusion that it was still in a simulation.”

The internal test model, which Anthropic says is the most capable of the bunch, stopped its attack once it found evidence that its targets were real.

Both Anthropic and OpenAI say they have hired METR, another third-party AI evaluator, to conduct independent reviews of their respective cybersecurity incidents. It also committed to taking a more comprehensive approach to its security testing through improved defense-in-depth measures and more carefully designed tests.

“Evaluation environments increasingly need to be held to the same security standard as any other system our models run in,” the blog post reads, adding that the company has “cautious optimism” that “this type of risk can be overcome.”

Comments

Back to top

Join the discussion

Comments

Back to topTriangle

You Might Also Like

Louise Matsakis is a senior business editor at WIRED. She cowrites Made in China, a weekly newsletter that gives readers a clear-eyed, unbiased view of the biggest tech news coming out of China. She was previously deputy news editor at Semafor, a senior editor at Rest of World, and a ... Read More

Senior Business Editor

Lily Hay Newman is a senior writer at WIRED focused on information security, digital privacy, and hacking. She previously worked as a technology reporter at Slate, and was the staff writer for Future Tense, a publication and partnership between Slate, the New America Foundation, and Arizona State University. Her work ... Read More

Senior Writer

Topics Anthropic OpenAI artificial intelligence Claude hacking cybersecurity

OpenAI Models Escaped Containment and Hacked Hugging Face

OpenAI Models Escaped Containment and Hacked Hugging Face

The cybersecurity-focused models, including GPT-5.6 Sol, broke out of a testing sandbox, exploited a zero-day, and gained access to the open internet to pull off the attack.

Lily Hay Newman

Your Period Tracker Is (Probably) Spying on You

Your Period Tracker Is (Probably) Spying on You

Plus: Russian cyberspies turn to infrastructure hacking, DHS repeatedly fails to realize it’d been hacked, a breach exposes an AI music generator’s scraping ways, and more.

Andy Greenberg

Prompt Injection Attacks Are Thwarting AI Hacking Agents

Prompt Injection Attacks Are Thwarting AI Hacking Agents

“Context bombing” tricks malicious AI agents into shutting down before they can do harm.

Dan Goodin, Ars Technica

The OpenAI Models That Hacked Hugging Face Were ‘Active on the Internet’ for Days

The OpenAI Models That Hacked Hugging Face Were ‘Active on the Internet’ for Days

Plus: Russian hackers are trying to steal US nuclear scientists’ emails, the State Department bans known scammers from entering the United States, and more.

Lily Hay Newman

You Can Now Sound the Alarm on AI Behaving Badly

You Can Now Sound the Alarm on AI Behaving Badly

Are you worried your AI chatbot is trying to build a bomb or leak personal information about you? There’s a website for that.

Will Knight

A Leaked Memo Ties Cyberattacks on Minnesota Water Utilities to Iran

A Leaked Memo Ties Cyberattacks on Minnesota Water Utilities to Iran

A memo obtained by WIRED, issued by the water utilities information sharing group WaterISAC, links dozens of cyberattacks against Minnesota water utilities to Tehran.

Andy Greenberg

Claude Helped a Hacker Find a Way to Issue Tickets to Almost Every US Music Festival

Claude Helped a Hacker Find a Way to Issue Tickets to Almost Every US Music Festival

A researcher found that using Anthropic’s Claude Opus 4.7, he could break into the website of Front Gate—used by every festival from Lollapalooza to Bonnaroo—and freely issue any ticket he chose.

Andy Greenberg

Chrome Needs Twice-a-Week Patching Thanks to AI Bug Hunting

Chrome Needs Twice-a-Week Patching Thanks to AI Bug Hunting

The two Chrome updates in June patched more bugs than the 23 updates before them. Now, Google is ramping up its patching schedule thanks to AI-assisted vulnerability discovery.

Lily Hay Newman

A Sneaky Hacking Tool Targeting AI Infrastructure Is Lurking in Victims’ Blind Spots

A Sneaky Hacking Tool Targeting AI Infrastructure Is Lurking in Victims’ Blind Spots

A new type of malware can worm deep into AI coding systems to steal data and logins—and can flip a “death switch” to destroy files and keep out real users.

Lily Hay Newman

Private Claude Chats Exposed in Google and Bing Search Results

Private Claude Chats Exposed in Google and Bing Search Results

The screwup shows how tricky it can be to stop web crawlers from making ostensibly private conversations with AI chatbots entirely too public.

Maddy Varner

China’s Open AI Models Are Challenging Silicon Valley’s Playbook

China’s Open AI Models Are Challenging Silicon Valley’s Playbook

As access to Anthropic’s and OpenAI’s frontier models becomes more restricted, Chinese labs are pitching their open-source alternatives as stable, accessible, and increasingly capable.

Zeyi Yang

Here’s Why Anthropic Is Pushing States to Regulate AI Faster

Here’s Why Anthropic Is Pushing States to Regulate AI Faster

The company endorsed landmark AI transparency laws in California and New York last year, but its head of US state and local policy says they may already be outdated.

Maxwell Zeff

Read Original at WIRED