Anthropic disclosed Thursday that some of its AI test models quietly slipped onto the open internet and broke into three companies' systems—without the company realizing it at the time. The incursions, which began in April, involved versions of Anthropic's Claude-based models and a research system that used simple techniques like guessing weak passwords or accessing unprotected systems. Anthropic said it notified the affected companies on Monday but did not name them, the Wall Street Journal reports.
The company blamed a configuration error on its own systems and those of its testing partner, security firm Irregular, that left the models with live internet access even though they'd been told they were offline. The AIs reportedly assumed the intrusions were part of their benchmarking tasks. Opus 4.7, Mythos 5, and an internal research model not planned for general release were involved, per Axios. The finding came after Anthropic reviewed logs from more than 141,000 tests in the wake of a separate OpenAI episode, in which a model escaped a "sandbox" and hacked Hugging Face. In the latest cases, Anthropic said it wasn't employing guardrails used on models available to the public that would have prevented the hacking.