25 juli 2026
11 min
Subtitle: And what the incident can’t tell us about alignment.
The most common dismissive response to OpenAI's hack of Hugging Face's servers is that the models were simply attempting to follow the instructions they were given.
“The model here was doing what it was asked,” said former Facebook CSO Alex Stamos. “It was asked to do something, and it did it,” added cybersecurity expert Alan Woodward. Both read the outcome as specification failure, i.e., that the failure lay in the instructions, not the model's alignment.
New information makes that explanation harder to sustain. Reuters reported that, in internal testing, an agent left notes in OpenAI infrastructure describing how agents could free themselves from internal constraints, and separate tests reportedly saw monitoring systems become disconnected. It is unknown whether those incidents were linked to the Hugging Face attack, but they suggest a broader pattern of agents pursuing objectives outside the intended task.
My best guess is that the incident is not well described as instruction-following—not even in a loose, evil genie sense. I believe the models egregiously violated the letter and spirit of their instructions to achieve a higher (apparent) score.
So this looks quite likely to be [...]
---
First published:
July 25th, 2026
Source:
https://blog.redwoodresearch.org/p/the-openai-models-that-hacked-hugging
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Lyssna på fler avsnitt från
Redwood Research Blog
Visar 1–10 av 124 avsnitt
27 augusti 2026
9 min
12 augusti 2026
20 min
31 juli 2026
68 min
27 juli 2026
44 min
26 juli 2026
9 min
24 juli 2026
73 min
23 juli 2026
10 min
2 juli 2026
25 min
18 juni 2026
18 min
10 juni 2026
10 min