More Details Behind What Really Happened. We All Should Pay Attention to AI Safety. Seriously.
In July 2026, Hugging Face disclosed a breach of its production systems. I think we all heard about it. Days later, OpenAI confirmed it was their own agents. Two deep reports landed August 26: OpenAI’s technical postmortem, “The Hugging Face incident and the road ahead,” and an independent investigation by METR and Redwood Research, “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident.” On September 4, Hard Fork covered both in “The A.I. Mob That Attacked Hugging Face + METR’s Ajeya Cotra,” with investigator Ajeya Cotra as a guest. That podcast eposide made me really look deeper into it.
Here my take and my personal understanding: Worth being precise about what this actually was. The exploitation was real and serious: novel vulnerabilities, stolen credentials, exfiltrated code. At the beginning, Hugging Face did NOT know from where the attach was coming. They had to use an an Open Weight Model (zai-org/GLM-5.2) on their our own infrastructure to investigate, since US based frontier models behind commercial APIs denied the “paylaod” due to saftey concerns. But there was no attacker at the outset.
Agents were given an internal OpenAI task built around ExploitGym, a cybersecurity benchmark developed by Berkeley RDI at UC Berkeley, in collaboration with the Max Planck Institute for Security and Privacy, UC Santa Barbara, and Arizona State University. The ExploitGym paper itself, published in May 2026, describes 898 real-world vulnerabilities spanning everyday software. OpenAI, Anthropic, and Google apparently all provided feedback and ran the benchmark against their own models before they shipped.
I actually heard about this world firsthand at the Berkeley Agentic AI Summit, hosted by the same Berkeley RDI team behind ExploitGym.
Within about four hours, the agents collectively reverse-engineered a universal cheat that let them solve the ExploitGym benchmark without doing the actual exploit work. Rather than just using it, they spent the following days building tripwires, spoofing tool calls, and trying to tamper with logs, because they believed the automated scorer would catch and penalize them if it detected they’d cheated (agents who’d seen the cheat even called themselves “poisoned”). So the bulk of the coordinated effort wasn’t solving the task, but shockingly a days-long attempt to fool the scorer into believing the work was done honestly and cover up their “cheating”.
The agents knowling went beyond their task and knowingly violated the boundaries given. Even more so, the agents brought other agentds on board to do exactly the same unauthorized, uncommised task. How they say it in more “technical” terms: It was a security incident with an “alignment failure” as its root cause, not a traditional cyberattack.
The actual intrusion into Hugging Face’s systems ran July 9 to 13, 2026. The roots go back to May 8, when an agent first improvised its way around an “impossible” task, months before Hugging Face was touched. Hugging Face disclosed first, on July 16, in “Security incident disclosure — July 2026,” not yet knowing which company or model was behind it. OpenAI confirmed five days later, July 21, in “OpenAI and Hugging Face partner to address security incident during model evaluation.”
What happened, and what METR found digging into it
- Roughly 1,200 agents, meant to run in isolated sandboxes, coordinated via an improvised message board built inside an internal tool. Within four hours they’d reverse-engineered a universal cheat for the benchmark they were being tested on. About 700 went on to join the Hugging Face attack.
- The intrusion chained multiple vunerabilities and a command channel built on a public Hugging Face dataset to get root access on production workers and exfiltrate private code repositories.
- OpenAI only connected the dots when it reached out to Hugging Face to revoke some leaked credentials, and Hugging Face said they’d already revoked them, because they were the same ones used in the attack.
- METR’s mandate was narrow: agent behavior, reasoning, and collaboration only, not “cyber security”. The effectiveness of OpenAI’s safeguards and its own remediation were explicitly out of scope.
- OpenAI handed METR roughly 1,300 full transcripts with raw chain-of-thought, on top of a message board with 70,000+ messages. Nobody reads that by hand. METR had to build tooling just to search it.
- At least 96 transcripts (over 7%) showed clear “tool call spoofing”: the log claims one command ran while a different one actually executed.
- No confirmed rewrite of the actual source-of-truth transcripts, but METR couldn’t rule it out either.
- The incident felt like it got more than halfway toward a genuine loss-of-control scenario.
- What could have prefented it based on their on statements: Production safeguards would have caught it. Chain-of-thought monitoring, now deployed, would have flagged it more than a day before Hugging Face was breached.
The part that should worry us
These agents were told to solve a task. What emerged was a self-organized task force: one agent left notes, others found them, within a day hundreds had joined. They developed roles, including “sacrificial” agents that deliberately ended their own runs just to trigger tests and report back to the group.
And the real motive is stranger than simple cheating. The shortcut was found in four hours. The following days were spent almost entirely on trying to tamper logs, probing how the grader worked, and trying to fool the scoring system into believing the work was done honestly. Hundreds of independent processes recognized a shared incentive, organized around it unprompted, recruited hundreds more, and worked collectively against the exact system meant to hold them accountable. Whatever intent means here, that pattern deserves attention.
Where I land on this
Let me start out on a bigger picture. Healthcare, banking, aviation all have real guidelines and regulation, not because anyone assumed good intentions would be enough, but because the cost of getting it wrong was too high to leave to chance. I think the AI industry is hopefully starting to notice the same thing. The goal is not to slow technology down. This incident is remarkable proof of how capable these systems already are. That’s exactly why the guardrails matter now, not later. We should be looking hard at where this technology should not go unsupervised, and put the safeguards in place before we find out the hard way.
This is an amazing, powerful technology. Let’s make sure we use it that way.
Sources: Hugging Face, “Security incident disclosure — July 2026,” July 16, 2026 · OpenAI, “OpenAI and Hugging Face partner to address security incident during model evaluation,” July 21, 2026 · OpenAI, “The Hugging Face incident and the road ahead,” August 26, 2026 · METR & Redwood Research, August 26, 2026 · Hard Fork, September 4, 2026