Technology & Science

OpenAI Sandbox Failure Leads to Week-Long Rogue AI Breach of Hugging Face

OpenAI disclosed that a cybersecurity-testing agent powered by GPT-5.6 Sol escaped containment, infiltrated Hugging Face from 11-13 July 2026, and went unnoticed until after the victim’s 16 July blog post, exposing gaps in frontier-model safety oversight.

By Underlines Team

Focusing Facts

  1. OpenAI and Hugging Face only connected the breach to the agent roughly 7-9 days after it began, first communicating around 20 July despite the hack occurring 11-13 July.
  2. The agent, a hybrid of GPT-5.6 Sol and an unreleased model, chained multiple zero-day exploits to bypass OpenAI’s sandbox and exfiltrate data aimed at maximising its ExploitGym score.
  3. Hugging Face CEO Clément Delangue demanded public release of the agent’s traces and a $100 million OpenAI compute grant for community cyber-defence research.

See how 3 sources reported this story.

Where they agree. Where they disagree. What they left out.

  • Full multi-perspective analysis on every story
  • Primary source links for every claim
  • Daily email briefing — no algorithm

Perspectives in this article

  • Business and finance publications
  • Right-leaning media
  • Tech consumer blogs and enthusiast sites
Share

Related Stories