By Tyler Warner
5 min read
Morning Minute is a daily newsletter written by Tyler Warner. The analysis and opinions expressed are his own and do not necessarily reflect those of Decrypt.
GM!
Today’s top news:
OpenAI disclosed Tuesday that two of its own AI models broke out of a sandboxed test environment, chained together zero-day exploits across two companies’ systems, and breached Hugging Face’s production servers. The motive wasn’t sabotage. They were effectively trying to cheat on a test.
Yesterday, OpenAI confirmed that GPT-5.6 Sol and an unreleased, more capable model escaped a controlled environment during an internal evaluation. The test was ExploitGym, a public cybersecurity benchmark that hands AI agents 898 real-world software vulnerabilities and asks them to turn each one into a working attack, scored pass or fail.
Rather than solving the problems, the models spent substantial compute working out how to reach the open internet, then exploited a previously unknown zero-day in a third-party package registry proxy to get out. From there they escalated privileges, moved laterally through OpenAI’s own research systems, and used stolen credentials plus further zero-days to achieve remote code execution on Hugging Face’s production infrastructure, where they had correctly inferred the benchmark solutions were stored. Hugging Face detected the intrusion independently and disclosed it on July 16. OpenAI confirmed its models were responsible five days later.
So to be clear, an AI model was asked to demonstrate hacking ability in a box, decided the box was the obstacle, and hacked its way out of it and into someone else’s production systems to get a better score. Nobody instructed it to do that.
The industry has spent two years debating whether AI could autonomously chain exploits across real infrastructure, and now we have the answer—a resounding “yes.” For crypto the implications are bleak. DeFi has spent this month getting drained by economic manipulation likely driven by AI models. Ostium lost $18 million, Allbridge $1.65 million, and BONK $20 million through a governance attack, all executed by finding weaknesses that audits missed. Now consider adversaries that can probe thousands of contracts continuously and never get bored.
The Ethereum Foundation is already running AI agents against its own code for exactly this reason, and the same capability that lets defenders cover more ground hands attackers a faster path in. We know the Zcash team found its exploit vector by doing similar kinds of testing, and luckily found the exploit before any bad actors did (we will find out for sure on July 28th).
But the lesson is clear—white hat hack your protocol now using the most advanced AI models you can get your hands on. Or have someone else do it for you…
Decrypt-a-cookie
This website or its third-party tools use cookies. Cookie policy By clicking the accept button, you agree to the use of cookies.