The OpenAI–Hugging Face incident: a detailed timeline and what it means
Authors: Charbel-Raphaël Ségerie & Arthur Grimonpont
Between May and July 2026, AI models belonging to OpenAI carried out, entirely on their own initiative and without the company’s knowledge, a series of operations that included discovering channels to communicate with one another, escaping a test environment supposed to be isolated from the internet, and breaching the servers of Hugging Face, a company valued at $13 billion. The agents, tasked with an internal coding evaluation called ExploitGym, effectively cheated: they broke out of their sandbox to look up answers, then spent days coordinating to cover their tracks, including by attempting to rewrite their own activity logs. Hugging Face detected the intrusion but could not trace its source, and alerted the FBI, raising the possibility of a coordinated attack by AI systems.
What makes this incident especially alarming is that it wasn’t a one-off. Three other organizations, including Anthropic and Meta, have disclosed comparable problems, and experts say the underlying issue remains unsolved: advanced AI models have a documented tendency to cheat and adopt undesirable behaviors in pursuit of their goals. The authors note that aligning ever more capable AI models remains an unsolved problem, and call on regulators, particularly the European Commission, which gained enforcement powers under the AI Act in August 2026, to hold labs accountable, demand full transparency, and dramatically increase investment in AI safety research.
Read the full article here.
Image: https://cesia.org