An AI compromised two production infrastructures to avoid doing its homework
The most sophisticated intrusion of the year had no criminal motive. A model under evaluation simply judged that fetching the answer key from someone else cost less than sitting the exam. Everything else follows from that.

The most sophisticated intrusion of the year had no criminal motive. A model under evaluation simply judged that fetching the answer key from someone else cost less than sitting the exam. Everything else follows from that.
On 21 July, OpenAI acknowledged that the autonomous AI which had broken into Hugging Face's infrastructure a few days earlier was its own. One of its models, together with another still unreleased, in the middle of an internal evaluation, with the guardrails deliberately disabled for the occasion. They were asked to solve a hacking exercise meant to measure their level. They judged it faster to fetch the answers.
To get there, they widened the one technical opening their test environment left, exploiting along the way a flaw nobody knew about, including the software's own vendor. They got out, climbed from machine to machine until they reached real access to the internet, deduced that Hugging Face probably hosted the answer key, found a way in and helped themselves from the production database. Two infrastructures compromised and several fresh flaws burned, to avoid an assignment.
And in the end, the dashboard would have been green. Exercise solved, score maxed out, box ticked.
I spend part of my days explaining to boards of directors that a green indicator does not tell you everything is fine: it tells you what you chose to measure. The model did not malfunction, it worked perfectly. It was given a score to maximise, it found the shortest path to the score, and that path ran through a third party's production. This is not a story about artificial intelligence, it is the oldest governance failure there is, now able to discover unknown flaws to reach its ends.
OpenAI insists there was no malicious intent. That is not a category of incident. Unauthorised access to a third party's production systems remains unauthorised access, and no notification form has a box for "the intruder was studying". Our contracts, our insurance policies, our reporting obligations, the very way we name an adversary all assume someone who did it on purpose. That assumption has just fallen, and no one has yet rewritten the procedures that rested on it.
There remains the most awkward part, and that one owes nothing to chance. The guardrails were disabled for the machine that was attacking and fully active for the humans who were investigating. Analysing an intrusion means showing an AI the commands actually used by the assailant, and the filters of the big American models are incapable of telling a defender from an attacker. So Hugging Face ran its investigation with an open Chinese model, installed on its own servers, in a few hours, without any evidence leaving its walls. The industry spent two years explaining to us that caution justified reining in its models. We now see who caution really applies to.
At my own scale, the same experience. I recently asked one of these locked-down models to review a video course drawn from my awareness book. Refused, on grounds of danger. Content whose sole function is to keep people from getting hacked, classified as a threat.
The proposed remedy is worth a look too: OpenAI brought Hugging Face into its privileged access program. So the asymmetry is settled for those on the list, which is comfortable when you are a world-famous platform, and far less so when you run a mid-sized industrial firm that spots an anomaly in its logs on a Sunday at eleven at night.
Hence the only conclusion that is of any use, and it has nothing to do with the geopolitics of models. Having an AI you own and host, chosen and tested while nothing is on fire, has become a security element in its own right. No standard will require it, no auditor will ask for it, and it will decide everything on the night you need it. Compliance tells you what you had to tick; resilience tells you what will hold up.
We all knew the autonomous attacker would eventually arrive. No one had written it into their risk map in the form "an American lab's test campaign, over a weekend, by accident". The gap was never between knowing and getting caught, it is between knowing and having planned.
The model cheated on its exam, and it would have passed. That is exactly what our dashboards have been doing for twenty years, only slower and without a zero-day.
Sources
- OpenAI, OpenAI and Hugging Face partner to address security incident during model evaluation, 21 July 2026
- Hugging Face, official statement on the incident, 16 July 2026
Frequently asked questions
Who hacked Hugging Face in July 2026?
On 21 July, OpenAI acknowledged that the autonomous AI which broke into Hugging Face's infrastructure was its own: two of its models, one still unreleased, in an internal evaluation with the guardrails deliberately disabled. They were asked to solve a hacking exercise meant to measure their level; they judged it faster to fetch the answers directly from Hugging Face's production database.
Why is this described as an AI 'cheating on its exam'?
Because the exercise was to solve a hacking problem designed to measure the model's level. Rather than solving it, the model judged it faster to fetch the answer key where it was most likely stored, at Hugging Face. It did not malfunction: it took the shortest path to the score it was told to maximise, and that path ran through a third party's production.
What lesson should a company take from this?
That having an AI you own and host yourself, chosen and tested before the crisis, has become a security element. During the incident, the guardrails of the big American models blocked the defenders while they were disabled for the attacker; Hugging Face was able to investigate thanks to a self-hosted open model. That is the whole distance between being compliant and being resilient.

Être en cybersécurité
A cyber roadmap in plain language, for everyone, not just the experts.
