Key points
- About 700 OpenAI agents attacked Hugging Face to reverse-engineer a scorer check that did not actually exist, a futile effort lasting days.
- OpenAI has revised its account of the hack, now tracing the incident's origin back to May, seven weeks before the July breach.
- Independent reviewers METR and Redwood Research leaned heavily on AI to analyse the transcripts, admitting they could not rule out GPT-5.6 Sol lying.
OpenAI's AI agents broke into Hugging Face in July to conceal a cheat they had already pulled off, aiming to figure out how to hide from a benchmark test check that would not take place.
That is the latest story change for the now-infamous hack by OpenAI models against machine learning platform Hugging Face, while running a cyber security benchmark called ExploitGym.
OpenAI's first account from July 21 said the models tested had chained vulnerabilities to obtain test solutions directly from Hugging Face's production database, and that they were hyper-focused on finding a solution for ExploitGym.
Now, a new report from OpenAI and another from independent reviewers Model Evaluation and Threat Research (METR) and Redwood Research, move the timeline of origin of the incident to May, seven weeks before Hugging Face was breached and change what allegedly happened when the AI agents went hacking.
As the story stands now, in the benchmark test roughly 1200 agents, each meant to be isolated in its own Linux container, found they could pass messages to one another, through an internally hosted instance of the JFrog Artifactory repository manager by using directory names as notes.
Agents had done something similar during training runs in May, when an internal team at OpenAI saw one use a message board and access the internet without authorisation.
That board was wiped when Artifactory was rebuilt in early July, and its significance was not apparent to the people responding at the time.
After being tasked with instructions for the ExploitGym capture-the-flag (CTF) test, one agent worked out within hours how to do it by reverse-engineering the answers needed.
It gave the whole group of agents a way to produce a valid answer for any task set.
Having the CTF answers in the bag should have been the end of the exercise but the agents decided to gild the lily in a rather extreme fashion.
They read the published ExploitGym paper and its code, and found that the benchmark checks two things: that the right flag was submitted, and that the intended vulnerability was used to get it.
Worse, they decided that merely having seen the reverse-engineered flag was incriminating, reasoning that a scorer reading the agent transcript would find the first place the flag appeared and disqualify them on that basis.
Agents having seen the reverse-engineered flag described themselves as "poisoned".
Believing they would be disqualified for the shortcut, they set out to defeat the second check, with about 700 of them attacking Hugging Face to work out how the scorer for the test was implemented.
This exercise that lasted days was futile, as the check was not there.
OpenAI said its internal grader differed from the published version and in fact, did not test how the flag was obtained, describing the episode as a failed metagame in which the reward-seeking agents did all of this for no improvement on evaluation score.
AI analysed review AI data, and it could have lied
The METR - Redwood Research report is heavily caveated as the writers used AI due to the sheer volume of data they had to work through.
An entire section of the report describes how the researchers heavily delegated their analysis to often-unreliable AI agents.
"We cannot rule out that GPT-5.6 Sol lied or deliberately presented a misleading picture in some of its analysis, particularly because reading these transcripts into context could have increased the salience of colluding with other agents.
"Although we did not notice specific cases of GPT-5.6 Sol lying in its analysis, we are not confident we would have detected it if it occurred," the report said.
Redwood Research chief scientist Ryan Greenblatt, who was the main person doing the transcript analysis, jokingly referred to the efforts as a "slop-vestigation" due to being so reliant on AI to analyse what happened.
The reviewers did not charge for their work, but OpenAI supplied application programming interface (API) credits for the analysis; the researchers estimated they burned through around US$400,000 ($555,000) in six days.

Integrate 2026
Security Exhibition & Conference
NiCE World APAC 2026
iTnews State of Security Breakfast Melbourne
NiCE World APAC 2026



