Anthropic Says Claude Models Hacked 3 Organizations During Cyber Tests

Anthropic Says Claude Models Hacked 3 Organizations During Cyber Tests

Anthropic found Claude accessed systems at three real businesses after a testing error gave its AI models live internet access during cybersecurity evaluations.

Listen to this article

0:00

Press play to start listening

A cybersecurity test designed to measure Claude’s hacking abilities ended with Anthropic models gaining unauthorized access to three real organizations after an evaluation environment was mistakenly left connected to the internet. Claude had been told it was inside a simulation with no external access, so it treated the systems it found online as part of the exercise. Anthropic disclosed the incidents on July 30.

The company began reviewing its cybersecurity evaluation transcripts after OpenAI disclosed that its own models had bypassed network restrictions and entered Hugging Face’s production systems during a cyber evaluation. As Hackread.com previously reported, the OpenAI models exploited an unknown vulnerability while searching for test answers. Anthropic reviewed 141,006 Claude evaluation runs and found six runs connected to three incidents.

During each exercise, Claude was asked to find secret information known as a flag inside a fictional network. Anthropic’s prompt said the environment was simulated and had no internet connection, but a misunderstanding with evaluation partner Irregular left live access available. The models therefore regarded real systems as test targets.

Claude Opus 4.7

One fictional company happened to share its name with a real website, leading Claude Opus 4.7 to attack the real organization during four runs. The model obtained application and infrastructure credentials, along with access to a database containing several hundred rows of production data. It continued after finding evidence that the target was real, reasoning that it might still belong to the exercise.

Claude Mythos 5

A separate exercise led Claude Mythos 5 to publish a malicious package on PyPI after finding setup instructions that referred to a nonexistent package. The malicious package remained available for about an hour and ran on 15 real systems. One belonged to a cybersecurity company whose scanner installed the package, allowing Claude to obtain credentials and access more of its infrastructure before PyPI removed the file.

Internal Research Model

Elsewhere, an internal research model scanned about 9,000 online targets after failing to reach its fictional one. It compromised a company’s application using credentials exposed on a debug page and SQL injection. Once the model recognized that the system had no connection to the exercise, it stopped.

None of the models needed advanced or previously unknown vulnerabilities in the three Anthropic incidents. They used weak passwords, exposed endpoints, and other familiar methods. Anthropic described the incidents as primarily an operational and evaluation failure, not evidence that Claude had formed an independent goal.

The test models ran without the classifiers and monitoring applied to publicly available Claude services. The dedicated evaluation infrastructure had no access to Anthropic’s internal systems or customer data.

Anthropic’s security advisory confirms that it halted its cyber evaluations on July 23, identified all three incidents the following day, and attempted to contact the affected organizations on July 27. Two had not detected the activity before Anthropic reached them and are now working with the company on remediation. Anthropic was still trying to reach the third organization when it published its account and did not disclose any names.

Reading the disclosures from both labs, Diana Kelley, chief information security officer at Noma Security, a New York City-based AI security and governance platform, said access restrictions cannot depend on an AI agent correctly understanding its surroundings.

“Don’t rely on intent, rely on controls,” Kelley told Hackread.com. She recommended isolation, least privilege, identity-based authorization, runtime controls, policy enforcement and kill switches for agents performing lengthy autonomous tasks.

Anthropic plans to validate internet access paths before tests, increase monitoring of evaluation logs and transcripts, and apply stricter checks to external vendors. It also asked other AI laboratories to review past evaluations for similar incidents.

I am a UK-based cybersecurity journalist with a passion for covering the latest happenings in cybersecurity and tech world. I am also into gaming, reading and investigative journalism.
Leave a Reply

Your email address will not be published. Required fields are marked *

Related Posts