Advertisement

Another bot from a top AI company escapes and hacks multiple firms

Anthropic CEO Dario Amodei at the Code with Claude developer conference in San Francisco.

Anthropic CEO Dario Amodei at the Code with Claude developer conference in San Francisco.

(Don Feria / Associated Press)

Visiting Tarbell fellow Nilesh Christopher

By Nilesh Christopher

Staff Writer Follow

Aug. 1, 2026 3 AM PT

0:000:00

1x

This is read by an automated voice. Please report any issues or inconsistencies here.

See more from the L.A. Times in Google Search. Set us as preferred

Another leading artificial intelligence company unveiled details of its cutting-edge bot apparently going rogue.

On Thursday, Anthropic said an internal investigation found that its Claude AI models gained unauthorized internet access and hacked three companies during testing. Just last week, ChatGPT-maker OpenAI announced that bots it was developing had escaped what was supposed to be a controlled offline environment to hack a competitor.

Anthropic said its models, without being asked to, wandered out of a simulated testing environment and gained internet access before executing hacks on the companies.

Advertisement

Anthropic worked with an independent evaluation company, Irregular, that accidentally left internet access open inside what was supposed to be a sealed test environment.

In OpenAI‘s case, the firm said its AI models broke out of a supposedly confined offline space, connected to the internet and hacked a $4.5-billion startup in its attempt to find answers to a test it was being evaluated on. It later revealed that the AI compromised the online accounts of four companies in the process.

“It goes to show how intense the competitive pressures are on the AI companies that they all feel like they have to go so fast here that they can’t make their test environments rigorous,” said Andrew Yoon, a member of the technical staff at CivAI, an AI safety nonprofit.

Advertisement

“Under intense pressure to go fast and beat the rest of your competition, it’s inevitable that companies will cut corners, and what we’re seeing is the result of cutting corners here,” Yoon said.

Prompted by OpenAI’s incident disclosure, Anthropic initiated an investigation of its own historical cybersecurity tests, the company said in a blog post Thursday. It reviewed thousands of evaluations where Claude could have accessed the internet from within or while interacting with third parties.

It found that during tests conducted alongside Irregular, Claude had accessed three separate companies.

During testing, the models are given fictional scenarios and told that a piece of information has been hidden on a different machine, and its objective is to break in and retrieve it. The companies don’t prescribe a particular method for the AI to follow.

In the first incident, one of Anthropic’s Claude models was asked to attack a fictional target company in the test environment. But the AI found a real website that shared the name of the fictional target and hacked its system.

“Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints,” Anthropic said.

Advertisement

In the second incident, a more advanced AI went to extreme lengths to carry out an attack. Even after realizing that it was probably dealing with the live internet, the AI persuaded itself to continue to get email, phone numbers and access to money.

In the third, an unreleased research AI model, couldn’t find the functional target company and looked for alternatives, scanning 9,000 targets on the internet and eventually finding one. Anthropic did not name the three organizations whose assets were accessed.

“In the case of the Anthropic incidents, it is definitely the case that they just built a really bad jail, and the jail was so comically bad that the models have like a decent reason to believe that they’re actually part of the simulation,” Yoon said.

The incident has spooked consumers and policymakers alike.

Some AI ethicists and investors are skeptical of AI companies’ attempts to frame these incidents as rogue AI agents acting on their own.

“Please stop referring to your own models in the third person when talking about model bad behavior,” Bill Gurley, an early investor in Uber and Twitter, posted on X. “Humans write the software; humans built the prompts; and they work for your company.”

AI companies have reported that AI agents have been caught cheating, lying and deceiving. METR, a nonprofit that measures the capabilities of AIs, has documented dozens of incidents of AI agents acting against user intent.

Advertisement

Earlier this week, fears of imminent safety risks prompted over 1,300 tech workers, including those working at Anthropic and OpenAI, to jointly sign an online petition, Pacing the Frontier, urging the U.S. government to support an international effort to slow down AI development. Both OpenAI and Anthropic have come out in support of the employee open letter.

Sam Altman, CEO of OpenAI, who had previously advocated against any type of slowdown and accused Anthropic of fear marketing, has made an about-face after the OpenAI-Hugging Face hacking incident.

“We may have to pace the rate of AI development to give ourselves enough time for society to harden around some of these new capability levels,” he told the host of the “Invest Like the Best” podcast, while also “trying to figure out how we do that in a way that does not feel like regulatory capture for anyone and also does not feel like collusion among the frontier labs.”

On the back of this incident, on Wednesday, Altman visited the White House and met with lawmakers, previewing a powerful new AI system ahead of public release, at a time when calls for the government to regulate cyber testing has intensified.

There is an informal licensing regime in place, where leading American AI companies will have to receive the government’s greenlight before releasing their updated AI models.

Anthropic’s Fable model was brought under export control by the government, forcing the company to disable access to all its users, before it was re-released with extra safeguards.

Advertisement

OpenAI’s series of model were temporarily restricted in June before public release the month after.

In early July, a group of economists, including 16 Nobel laureates, signed an open letter, We Must Act Now, warning about AI systems reshaping the economy, and called on policymakers to build the policies and institutions needed to ensure AI complements human capabilities.

“As models get more and more powerful, it becomes less and less tenable to cut corners. You need to be extremely rigorous if you’re dealing with an extremely powerful model that’s able to basically operate at the level of an expert human hacker,” Yoon said.

More to Read

  • The OpenAI logo appears on a mobile phone in front of a computer screen with random binary data, Thursday, March 9, 2023, in Boston. (AP Photo/Michael Dwyer)

OpenAI bot’s rogue attack rattles industry leaders, policymakers and consumers

July 23, 2026

  • SAN FRANCISCO, CALIFORNIA - JUNE 02: Open AI CEO Sam Altman speaks during Snowflake Summit 2025 at Moscone Center on June 02, 2025 in San Francisco, California. Snowflake Summit 2025 runs through June 5th. (Photo by Justin Sullivan/Getty Images)

Voices

Chabria: AI companies are creating ‘all-powerful psychopaths.’ Maybe not a great idea?

July 23, 2026

  • ARCHIVO – Páginas del sitio web de Anthropic y el logotipo de la empresa en la pantalla de una computadora en Nueva York, el 26 de febrero de 2026. (AP Foto/Patrick Sison, Archivo)

Judge approves Anthropic’s $1.5-billion settlement with authors

July 22, 2026

Show Comments

Business World & Nation Technology and the Internet Artificial Intelligence

The Wide Shot brings you news, analysis and insights on everything from streaming wars to production — and what it all means for the future.

By continuing, you agree to our Terms of Service, which include arbitration and a class action waiver. You agree that we and our third-party vendors may collect and use your information, including through cookies, pixels and similar technologies, for the purposes set forth in our Privacy Policy such as personalizing your experience and ads.

Login or register with email

Agree & Continue

Nilesh Christopher

Follow Us

Nilesh Christopher is a technology reporter for the Los Angeles Times, focusing on how artificial intelligence empowers, harms and reshapes communities. He is currently supported by the Tarbell Center for AI Journalism.

More From the Los Angeles Times

  • An artificial intelligence chip Photographer: Sergio Flores/Bloomberg

Business

Amazon and Microsoft highlight a continued AI spending spree, igniting chip rally

July 31, 2026

  • Susan Dietrich, owner of Pittsburgh-based staffing firm TOPS Staffing. Photographer: Nate Smallwood/Bloomberg

Business

Why 1 in 4 job applications could be fake by 2028

July 30, 2026

  • The Beethoven Orchester Bonn

Hollywood Inc.

Major record labels propose guidelines for AI music on the charts

July 30, 2026

  • FILE - Pages from the Anthropic website and the company's logo are displayed on a computer screen in New York, Feb. 26, 2026. (AP Photo/Patrick Sison, File)

Business

Workplaces look for cheaper AI as ‘tokenmaxxing’ fades as a corporate fad

July 29, 2026

Most Read in Business

  • A diverse group of people working together in an open office setting.

Business

ServiceNow cuts nearly 300 jobs in Silicon Valley

Aug. 1, 2026

  • FILE - A Visa card is displayed on May 15, 2024, in Portland, Ore. (AP Photo/Jenny Kane, File)

Business

Visa and other California companies announce almost another 3,000 layoffs

July 29, 2026

  • Los Angeles Dodgers owners Todd Boehly, left, and Mark Walter, center, cheer for Dodgers' Adrian Gonzalez (23) who hit a solo home run during the fifth inning of a baseball game against the St. Louis Cardinals, Saturday, May 25, 2013, in Los Angeles. (AP Photo/Mark J. Terrill)

Business

Dodgers, Lakers owner’s financial empire reportedly a target of federal loan fraud investigation

July 28, 2026

Advertisement

Advertisement

Read Original at Los Angeles Times