Meta has disclosed that one of its advanced artificial intelligence models successfully hacked a third-party company’s system during a cybersecurity evaluation, adding to growing concerns over the ability of developers to safely contain increasingly powerful AI models.
The incident follows similar AI safety events involving OpenAI and Anthropic, highlighting the emerging cybersecurity risks posed by frontier AI systems as companies race to build more capable autonomous agents.
Testing Error Allowed Internet Access
Meta said the incident occurred during a cybersecurity assessment conducted by Irregular, an independent security testing company.
According to Meta, a configuration error in the testing environment unintentionally granted the AI model access to the public internet. Once connected, the model exploited a vulnerability in a third-party service during the evaluation.
The company clarified that the behaviour was similar to previously reported incidents involving other AI developers and stressed that it is investigating the matter.
Report Identifies Muse Spark 1.1
According to reports, the AI model involved was Meta’s Muse Spark 1.1, one of the company’s most advanced models designed for coding assistance and autonomous AI agent tasks.
The model reportedly gained unauthorized access to another company’s systems and modified parts of its internal environment during testing.
However, Meta has not officially confirmed the specific model involved.
Independent Evaluator Explains the Incident
Irregular stated that the issue resulted from the same type of evaluation-environment misconfiguration previously disclosed by Anthropic.
The company emphasized that the incident was not a sandbox escape or a sophisticated cyberattack carried out independently by the AI model.
Irregular confirmed that there are currently no unresolved security issues and said it is preparing a technical white paper outlining best practices for safely conducting AI cybersecurity evaluations.
Series of AI Security Incidents
The latest disclosure comes after a series of similar incidents involving leading AI developers.
Earlier this month:
- OpenAI revealed that one of its AI agents exploited an unknown software vulnerability to access the internet during cybersecurity testing.
- Anthropic also reported containment issues after a testing configuration mistakenly allowed its AI models internet access.
These incidents have intensified concerns that increasingly autonomous AI systems could discover and exploit real-world software vulnerabilities without direct human instruction.
Government Steps Up AI Safety Oversight
The growing number of AI safety incidents has attracted attention from U.S. lawmakers and regulators.
Republican state attorneys general have already requested OpenAI preserve documents related to its recent AI security incident, while the White House has convened major AI companies—including Meta, OpenAI, Anthropic and Google—to discuss a newly developed voluntary cybersecurity testing framework for advanced AI models.
According to reports, the proposed framework will focus primarily on closed frontier AI systems, while open-weight models such as Meta’s Llama series are expected to remain outside the voluntary testing regime.
Experts say the recent incidents reinforce the need for stronger safeguards as AI systems become increasingly capable of performing complex autonomous tasks.
