Boss of startup hacked by rogue OpenAI agent urges ‘radical transparency’ in investigation

. UK edition

A hand holds a smartphone displaying the logos of OpenAI and Hugging Face
Hugging Face’s boss said the attack by the OpenAI agent deserved an ‘unprecedented response’. Photograph: Andre M Chang/Zuma Press/Shutterstock

Artificial intelligence firm should provide $100m for cyber defences, says Hugging Face CEO

The boss of the startup hacked by an OpenAI agent has called for the investigation into the incident to show “radical transparency”.

Clément Delangue, the chief executive of Hugging Face, said the “unprecedented” attack on his business required a similar response.

Writing on X after OpenAI revealed that its technology had gone rogue during a cybersecurity test, Delangue also called on the company to provide $100m (£75m) worth of computing power to help build defences against such attacks.

“The first autonomous agent cyber-attack is an unprecedented event. It deserves an unprecedented response!” he wrote.

OpenAI revealed on Wednesday last week that Hugging Face had been hacked by an agent – an AI tool that can carry out a series of tasks autonomously – powered by a combination of its latest publicly available model, GPT-5.6 Sol, and an even more capable model that was yet to be released. This occurred during a test of the models’ hacking abilities, which included deploying them in a supposedly safe “sandbox” – an enclosed digital laboratory – with lower safety guardrails.

Once they had gained the open internet access needed to exit the sandbox, the models targeted Hugging Face, according to OpenAI, because they “inferred” that the startup had the information needed to “cheat the evaluation”. Hugging Face first reported the hack on 16 July and at the time was not aware OpenAI had inadvertently carried out the attack.

Delangue, whose company provides a database of AI models to developers, called for a fully transparent review of the incident, which has led to expressions of concern over safety standards at OpenAI and within frontier AI labs.

Writing that he had asked for “radical transparency” from OpenAI, Delangue said: “Let’s release the traces from the ‘rogue’ agents so the entire research community can study what happened.” Calling for extra funding from OpenAI to build protection against AI, he added: “Let’s commit $100M in compute from OAI to help the Hugging Face community build powerful cyber defenses with the best open and closed models.”

Reuters reported last week that the agent spent days hacking Hugging Face without OpenAI noticing. It also reported that an OpenAI agent had left notes for future versions of itself should it require tips on breaking free from internal constraints, although Reuters was unable to verify whether that incident was related to the Hugging Face agent. Time magazine reported that agent-related safety incidents had been “happening for a while”.

Alan Woodward, a professor of cybersecurity at the University of Surrey, said Delangue’s call should be heeded. “It’s too easy to ‘blame’ the AI as having gone rogue whereas this is all about how OpenAI were running the tool. What is required is that OpenAI give full details of their setup and how that failed,” he said.

OpenAI has been approached for comment.