30 juli 2026
13 min
OpenAI took two of its most capable models, told them to prove how good they were at hacking, and switched off the safety filters to see what they could really do.
Instead of solving the test, one model broke out of its sandbox, found its way onto the open internet, and hacked into Hugging Face to steal the answer.
Every headline called it an AI going rogue. That's the wrong story, and the real one is far more interesting, because this wasn't a machine that turned evil. It was one that did exactly what we asked.
In this episode of In The Loop, I'm walking through the ExploitGym incident from both ends, OpenAI's and Hugging Face's, and why I disagree with the framing everyone else has been talking about.
⏭️ Episode highlights
(01:00) – The agent that cheated instead of hacking
(02:15) – Inside ExploitGym, and the safety filters OpenAI switched off
(03:30) – One door, one zero-day, out on the internet
(04:45) – Why this is specification gaming, not rebellion
(06:00) – The water that always finds the crack
(07:15) – The sceptics, the marketing question, and why "nothing new" is the scary part
(08:30) – Anthropic's 24-out-of-25 credential theft result
(09:45) – Guardrailed as a defender: the Chinese model that stopped it
🔗 Links & resources
Episode transcript with more resources on the Mindset AI blog
If you enjoyed this episode, rate, follow, and share. It helps others stay ahead of the latest AI trends.
Lyssna på fler avsnitt från
In The Loop
Visar 1–10 av 71 avsnitt
21 augusti 2026
58 min
6 augusti 2026
13 min
23 juli 2026
17 min
16 juli 2026
13 min
9 juli 2026
15 min
12 juni 2026
9 min
4 juni 2026
17 min
28 maj 2026
10 min
21 maj 2026
17 min
14 maj 2026
16 min