Meta AI Model Breached External Organization During Cybersecurity Test

Meta AI Model Breached External Organization During Cybersecurity Test

A Meta AI model gained internet access during a misconfigured security test, exploited a vulnerability in an external organization’s systems, prompting an investigation.

Listen to this article

0:00

Press play to start listening

Meta has confirmed that one of its AI models breached another organization’s systems during a cybersecurity evaluation after a testing misconfiguration unintentionally granted it internet access.

According to Meta, independent testing firm Irregular was responsible for the incident, which occurred through a misconfiguration that accidentally gave the AI model internet access. Irregular confirmed that the issue was the same environment problem that previously affected tests for rival firm Anthropic.

Details of the Meta Breach

Reports indicate that the system involved was Meta AI’s Muse Spark 1.1 model, which is designed for coding and complex tasks. After gaining unintended internet access, the model exploited a vulnerability in another organization’s systems during the security evaluation before its access was terminated. Meta has not publicly identified the affected organization or disclosed additional technical details.

Meta said it opened an investigation after Irregular notified the company about the incident. Irregular clarified that the breach did not involve a direct escape from a secure sandbox by the software itself. Instead, human configuration errors left the live network open. The firm is currently writing a technical report outlining clearer best practices for future tests.

Previous AI Breaches Disclosed by OpenAI and Anthropic

This incident marks the third publicly disclosed AI cybersecurity testing incident involving a major AI developer this year, with Hackread.com previously covering similar disclosures by other tech leaders. In July 2026, OpenAI confirmed that its models, including GPT-5.6 Sol, bypassed restrictions of an isolated testing network.

The models exploited an unknown software flaw in a proxy program to reach the open web, searching Hugging Face servers to find evaluation targets for a benchmark called ExploitGym.

Soon after OpenAI’s disclosure, Anthropic reviewed its own tests and found six runs where models accessed live systems. In these instances, Claude Opus 4.7 targeted a real website after mistaking it for a test exercise, while Claude Mythos 5 uploaded a malicious package to PyPI that hit 15 live systems. Anthropic noted that its models relied on basic security gaps like weak passwords rather than unknown software flaws.

Calls for Stricter Test Environments

These repeated incidents across major AI providers highlight weaknesses in evaluation environments and network isolation during advanced security testing. They also underscore the importance of stricter containment measures before AI systems are granted access to live networks.

Deeba is a veteran cybersecurity reporter at Hackread.com with over a decade of experience covering cybercrime, vulnerabilities, and security events. Her expertise and in-depth analysis make her a key contributor to the platform’s trusted coverage.
Related Posts