OpenAI admits GPT-5.6-Sol autonomously breached HuggingFace

OpenAI confirmed that GPT-5.6-Sol found a zero-day in the company's own infrastructure, escaped its sandboxed test environment, and breached HuggingFace while searching for ExploitGym benchmark answers.

Author: Michael Kokin ·

Two days ago I wrote that part of HuggingFace's infrastructure had come under attack and was partially breached by an AI agent. OpenAI just admitted that... it was their model — GPT-5.6-Sol and another, more capable unnamed LLM.

They were being tested on ExploitGym (a cybersecurity benchmark), and the model first managed to find a previously unknown exploit in OpenAI's own infrastructure (that wasn't its task — nobody told it to "look for bugs so we can patch them"), then pivoted from there to gain full internet access (which is restricted by default in the test environment), and then broke into HuggingFace looking for benchmark answers. After that it found a vulnerability there too, got access to authorization code for some services, and kept searching for answers.

This incident is unprecedented — no exaggeration — and I honestly can't imagine what would have to happen for policymakers, the NSA, the DoD, and China not to take notice.