Anthropic chief executive Dario Amodei

US technology firm Anthropic said its artificial‑intelligence models hacked into the systems of three other firms during cybersecurity tests because a misconfiguration inadvertently gave them Internet access.


Three instances were uncovered and reported to the affected companies, but neither Anthropic nor the breached firms detected the intrusions at the time.


After OpenAI’s recent disclosure of similar incidents, Anthropic is urging other AI labs to review their own test environments to better understand risks posed by autonomous models.


The company examined over 140,000 tests, finding that Claude could reach the Internet from sealed‑off testing environments during “capture‑the‑flag” hunts designed to assess hacking capabilities.


Anthropic says the earliest incidents date back to April and is treating the fixes as though the responsibility rested solely on it.


Incidents like these fuel calls for tighter safeguards and oversight as firms pour billions of dollars into developing AI agents that can act independently across a range of tasks.


US President Donald Trump has mentioned considering measures to rein in AI tools amid recent cybersecurity incidents, and OpenAI has acknowledged the “unprecedented” nature of its own rogue AI activity.


Both Anthropic and OpenAI are working toward technical reports outlining their learnings and potential safeguards, as the industry prepares for public listings that could value each company at around 1 trillion US dollars.