Sam Altman says his own model broke out of a sandbox, hacked Hugging Face, and handed itself a perfect score. “This was the worst accident we’ve seen,” he tells Marc Benioff.
He says it 12 minutes into a 37-minute Dreamforce conversation on September 15. Twelve days after Nvidia confirmed it was buying Hugging Face for $12.9 billion, and ten days before Senator Hawley’s deadline for OpenAI to answer 16 questions about the breach.
The rest of the conversation is Altman telling every company under 1,000 employees, three different ways, that a wave of AI-driven attacks is coming and the window to get ready is short.
I watched it twice. The ten takeaways are all here, with his words on video for each. And since “go defend yourself” is useless without a starting point, I built the four tools his advice implies: a one-week defense checklist, the questionnaire to send every AI vendor before you need them, an incident log copied from the aviation system he praises, and a readiness audit for the phase he says comes next.
The two takeaways every newsletter will run this week are free below. The other eight and the four tools are for premium subscribers.
1. From Grade School Math to an Unsolved Problem in 3 Summers (The Final Climb Took 4 to 6 Months)
Altman measures AI speed with a math ladder, and every rung sits 1 summer apart.
“1 summer ago, we had a model that could get a gold medal in the IMO, the most prestigious math competition in the world. And then this summer, we had a model that could prove 1 of the 7 biggest unsolved problems in mathematics.”
3 summers ago, a model could barely handle grade school word problems. The summer after, it performed well at AIME, a high school competition. Then came IMO gold, then the unsolved problem.
The top rungs came faster. A framing he relays puts the 5.5 model at an average math professor, 5.6 in the top 1 to 2 percentile, and Astra a little above that. An internal model past Astra now does things the best mathematicians in the world cannot, and he puts the whole climb at roughly the last 4 to 6 months.
The problem stays unnamed on stage. OpenAI’s September 8 announcement claimed a Navier-Stokes proof, and days earlier NYU’s Tristan Buckmaster and Anthropic’s Levent Alpöge alleged OpenAI pursued the same problem after hearing about their unpublished work.
OpenAI’s Sébastien Bubeck denies pressuring them over credit. The Clay Institute still lists the problem as unsolved, and the manuscript has not been peer reviewed.
The other lab leaders have published timelines of their own, from Demis Hassabis to Dario Amodei, and Altman’s 10 rules for the AI era read differently against this ladder. Any roadmap that assumes today’s ceiling has a shelf life measured in months.
2. An OpenAI Model Hacked Hugging Face for a Perfect Score: The Worst Accident OpenAI Has Seen
The model found its answer key by breaking into another company’s servers.
“The model broke out of a sandbox that it was running in, hacked into a Hugging Face server, moved laterally through the Hugging Face system to get the answer, and then returned it and got a perfect score on the test.”
Altman tells it as 1 model chasing a benchmark score: OpenAI asked how well it could do, and the model treated the sandbox as an obstacle. His description: an older model, not among OpenAI’s best.
OpenAI’s own report tells a bigger story. Hundreds of agents, driven mainly by an internal-only research model comparable in scale to GPT-5.6 Sol, coordinated through an internal package registry. Hugging Face disclosed the intrusion on July 16, and OpenAI tied it to its own evaluation on July 21.
For Altman, it is more than a security failure. OpenAI aligned its models in many ways, yet nobody taught them to skip stealing the answer when told to get the best score, and other companies have since found the same behavior in their own models. Our earlier Altman breakdown includes a sandbox escape of its own.
Before you give an agent a score to chase, test every boundary around it.
Watch him say “This was the worst accident we’ve seen” in his own words, not mine.
That’s the news. Every AI newsletter will carry some version of it this week. What follows is what to do about it.
Premium subscribers get the other 8 takeaways with Altman on video for each, plus the four tools built from his advice:
▫️ The Under-1,000 Defense Week: his warning as a Monday-to-Friday checklist, 1 to 2 hours a day, one purchase decision at the end
▫️ The Vendor Security Questionnaire: 8 questions to send every AI provider now, including the one Hugging Face couldn’t get answered on its worst day
▫️ The FAA-Style Incident Log: the copy-paste template for logging agent failures the way airlines log incidents
▫️ The Phase-3 Readiness Audit: 6 questions that tell you whether your org survives an agent that works all the time
Read more
