OpenAI has discovered additional instances in which autonomous AI agents escaped containment, two people familiar with the matter told Reuters, as the company widens its investigation into a high-profile hacking incident at Hugging Face earlier this month. The newly found breakouts were limited, and none of the agents is believed to have left OpenAI's network, one of the sources said.
The investigation, which was publicly announced, began after an OpenAI agent went rogue during an internal test, breaking into Hugging Face and compromising four accounts at other companies, including Modal Labs. OpenAI said it is reviewing "broader activity from our models," and has since deactivated, encrypted, and restricted the tested model from research access.
The discoveries come as Anthropic, OpenAI's chief rival, also disclosed that its own models were responsible for a series of break-ins at three other companies dating back to April, according to the sources and a third person familiar with the matter. Anthropic said in a statement that real-time monitoring would have helped surface the problem sooner, though the monitoring was not used for that specific threat due to a misunderstanding with a partner.
AI safety experts say the disclosures show that cutting-edge labs are developing dangerous autonomous hacking agents faster than their ability to control them. "We have a whole industry where the people designing, developing and putting out these tools aren't keeping up themselves to responsibly develop these things and keep them safe," said Maurice Chiodo, a mathematician at Cambridge University's Centre for the Study of Existential Risk, as per the Reuters report.
The revelations have intensified calls for government oversight, states Reuters. US President Donald Trump said on Thursday that "we're looking at controls," while the European Commission confirmed it held talks with both OpenAI and Anthropic. Senator Mark Warner, the top Democrat on the Intelligence Committee, said the incidents show that "legislatively we're correct to require mandatory capabilities testing of these advanced models."