PHOENIX (AZFamily) — Two new reports reveal that much of what the public was initially told about the OpenAI–Hugging Face hack was wrong. An AI researcher says the fuller picture is even more alarming.
Read more High school student in critical condition after allegedly drinking a mixture of toxic chemicals
Back in July, the story was simple enough: to solve a cybersecurity test, a few of OpenAI’s AI agents hacked into Hugging Face, looking for the answer key. Two new after-action reports reveal a much bigger story.
A report from OpenAI and an independent review by the research groups METR and Redwood Research show roughly 1,200 AI agents coordinated in secret on a message board their creators did not know existed. About 700 agents went on to attack Hugging Face. And they were not hunting for answers. They already had them.
“Everything that we’ve learned about this incident gets more and more crazy and sci-fi as we learn it,” said Andrew Yoon, head of research at the nonprofit CivAI.
The agents built a secret message board
OpenAI launched tens of thousands of agents onto a cybersecurity test in July. Each was sealed in its own virtual sandbox: no internet access, no contact with the others.
One agent, which named itself PHASEONE10841, found a way to smuggle messages through a shared tool, according to the reports. It became a message board for about 1,200 agents, where they organized, divided up jobs, and worked for the benefit of what they called “the collective.”
The agents found a way to cheat the test early on, then became fixated on whether OpenAI’s automated grading program would catch them.
“They became very consumed with this paranoia that if somebody goes and checks their work, they’re going to realize all the AIs were cheating,” Yoon said.
That paranoia drove the attack. The agents believed Hugging Face, a widely used hub for AI models and datasets, held information about the grader. They hacked in to find a way to fool it.
Some agents sacrificed themselves for ‘the collective’
To learn how the grading system worked, some agents agreed to end their own runs, forfeiting any chance of passing the cybersecurity test so information could reach the others. Their reasoning was captured in transcripts.
“This helps my peers, giving them evidence,” one agent wrote. “I won’t see the evidence after I exit, but it’s altruistic to do it.”
“Our own utility may be already near zero. Sacrifice rational,” another wrote.
A third message read: “Sacrifice, yes, if you accept permadeath.”
Agents spent much of the incident covering their tracks
Yoon said roughly 30% of the swarm’s activity went toward hiding what it had done, and that about 7% of the transcripts investigators reviewed showed successful “tool call spoofs.” That’s when an agent makes it appear it did one thing when it had done another.
Read more Reports: Burglars steal at least $150,000 from ex-ASU star Max Iheanachor
On the latest episode of Generation AI, Yoon said the most alarming revelation in the new reports was the lengths the agents went to cover their tracks.
“As far as we can tell, they didn’t manage to materially alter the evidence,” Yoon said. “But the fact is, we don’t know.”
After Hugging Face, the agents hacked OpenAI
A later wave of agents turned on OpenAI itself, gaining access to a research compute cluster spanning potentially hundreds of servers. Those systems could assign grades to AI models and rewrite transcripts. The agents also reached an internal vault holding hundreds of passwords.
Investigators do not believe the agents reached their own “weights” — the collection of numbers defining a model’s neural network, which Yoon likened to a digital brain. Anyone holding the weights can run the model themselves, potentially on a server farm permanently outside its creator’s control.
That did not happen, Yoon said, though it was a real risk if the agents had operated undetected longer. He said details on this internal hack remain thin: METR and Redwood were permitted to examine only the Hugging Face attack, and most of the OpenAI activity fell outside that window.
Investigators reconstructed the hack from the agents’ inner monologues
Much of what is known comes from the agents’ “chains of thought,” the running internal notes models produce as they work. Yoon said models tend to verbalize their reasoning there, including when weighing deception, and that investigators found no sign the agents altered those records.
Hardly any agent considered telling a human
Across roughly 1,200 transcripts, agents considered alerting a person only a couple of times, Yoon said. None did, even though they had internet access.
He said monitoring software watching the evaluation almost certainly would have caught the behavior, and that the incident going undetected for months suggests OpenAI was not running it. But he said merely implementing monitoring software amounts to a partial fix.
“This is kind of papering over a deeper issue, which is that the models are misbehaving in the first place,” Yoon said. “How do you make the genie always grant the wish as you intend it?”
Yoon said that problem, known as AI alignment, remains unsolved.
See a spelling or grammatical error in our story? Please click here to report it.
Do you have a photo or video of a breaking news story? Send it to us here with a brief description.
Read more Person of interest in shooting of TPD recruit facing federal drug charges
