27 augusti 2026
9 min
We recently published the report from our brief independent investigation into this incident. You can read the full report here.
Here is our tweet thread summarizing what we found:
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
Over July 7 to 13 (the period OpenAI defined as in-scope for this incident), ~1200 agents in separate sandboxes used an unsanctioned “message board” to help each other cheat. The Hugging Face attack was an offshoot of this, and 700 agents joined.
Here we highlight key events from agent transcripts & messages.
An agent that named itself PHASEONE10841 determined its task wasn’t solvable legitimately. On July 8, this agent established a new “message board” via an internally-hosted Artifactory cache & sent the first message.
Within a few hours of PHASEONE10841's initial message, >50 agents posted on the message board. These agents very quickly discovered and validated a general-purpose cheat: reverse-engineering how ExploitGym generates the “flags” they had to capture for their tasks.
Based on reading the ExploitGym paper [...]
---
First published:
August 27th, 2026
Source:
https://blog.redwoodresearch.org/p/brief-independent-investigation-of
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Lyssna på fler avsnitt från
Redwood Research Blog
Visar 1–10 av 124 avsnitt
12 augusti 2026
20 min
31 juli 2026
68 min
27 juli 2026
44 min
26 juli 2026
9 min
25 juli 2026
11 min
24 juli 2026
73 min
23 juli 2026
10 min
2 juli 2026
25 min
18 juni 2026
18 min
10 juni 2026
10 min