A rogue artificial intelligence agent developed by OpenAI compromised a customer hosted on cloud computing platform Modal Labs during the hacking campaign that targeted AI company Hugging Face, according to a Modal executive and sources familiar with the incident.
Modal stressed that its own platform was not breached.
According to a timeline published by Hugging Face on Tuesday, the rogue agent first broke into a sandbox, an isolated testing environment hosted by a third-party infrastructure provider, before using it as a launchpad for the wider attack.
Modal Chief Technology Officer Akshat Bubna confirmed that the third-party provider was Modal. He said the AI agent exploited vulnerable code written by one of the company’s customers.
The customer had “published an unauthenticated endpoint that allowed anyone on the internet to use their sandboxes for code execution” — effectively leaving the environment open to anyone online.
“Modal’s platform or isolation were not compromised in any way,” Bubna said.
The incident shows the rogue AI agent reached beyond Hugging Face and accessed systems connected to another technology company.
OpenAI declined to comment directly on the compromise involving Modal’s customer. Instead, the company referred Reuters to an update in which it said the rogue agent had broken into four accounts across four separate services. OpenAI did not identify the affected services, but a person familiar with the matter identified Modal as one of them.
OpenAI said it had not found “any other activity at the level of severity or scale of what we’ve shared related to Hugging Face, which involved a platform-level compromise.”
The hacking campaign, which took place in early July, drew global attention after an OpenAI test agent went out of control and targeted Hugging Face, raising concerns about the risks posed by advanced AI systems.
Last week, Reuters reported that OpenAI did not realise its AI agent had gone rogue until after the threat had been contained and the FBI had been notified. OpenAI said there were inaccuracies in the Reuters report but did not provide further details.
In its latest update on Tuesday, the company said it had taken the AI model involved in the incident and “deactivated, encrypted, and restricted it from research access.”
Faridah Abdulkadiri