Hugging Face, OpenAI, and Why You Should Care
So, here’s the story:
OpenAI was testing a model that it hasn’t yet released to the public.
Testing means a lot of different things. One of these things is to test how good a model is at hacking. Note that testing also means that normal cyber refusals had been reduced.
So OpenAI gave this unreleased model a machine (a computer), and asked this unreleased model to break into the machine. This is a deliberate simplification, by the way, you can read this paper to get a sense of what was actually going on.
The model wasn’t given internet access, so it was supposed to figure out how to hack into the machine on its own.
Except it couldn’t, even after lots of trying.
Now, the thing with these new models is that they just don’t know when to quit. We have, in fact, trained them not to know when to quit, and that’s one reason why they are so good at so many different things. It is because they keep trying different things.
So, this particular model, when it eventually failed at the task that it had been given, it also asked itself if it could try a different thing.
The model, remember, had not been given access to the internet. But it did have access to a software package manager - a small program that allows one to install many different software packages.
So what the model did is that it figured out how to exploit a vulnerability in the software package manager. This is the first of a series of Very Big Effing Deals.
Now, Hugging Face.1 Hugging Face is a bit like a shared Google Drive, but for AI models.
Once that unreleased model exploited the software package manager, it then used that software package manager’s ability to access the internet, and then it hacked Hugging Face’s servers. This is the second of that series of Very Big Effing Deals. It did this because it figured out that Hugging Face’s servers were likely to have answers or clues for the problem that it had been tasked with solving.
The model, as Alex Tabarrok guesses, was “on the loose” for at least a matter of days, probably around a week. “On the loose” means exactly what you think it means: it was able to use the internet, undetected, for about a week. This is the third of that series of Very Big Effing Deals.
Hugging Face knew it was being attacked, but it had no clue who was attacking it. They tried to use the best models at their disposal for defense, but those best models thought these requests were a trick, and refused to play ball.
So Hugging Face had to use a Chinese model (GLM 5.2) to investigate what was being done to its servers by an unreleased American model. This is the fourth of that series of Very Big Effing Deals.
Here’s the official response, but this is a more relatable take. And for dessert, see this, and think about what the state of play is likely to be a year from now.
Why do I say life is about to get exciting? Because we need to ask ourselves what models like these will do when released to the public and what will happen when even more powerful models are available in three months, six months, nine months from now. What about a year from now? What about financial systems? Electricity grids? National defense systems? Airports and other transport hubs? How good do you think their online security systems are?
And that is why life is about to get exciting.
On an entirely related note, the latest episode of AiVey, the podcast that I and Navin Kabra have started, features Rohit Sriwastwa, a guy who knows a thing or thirty about being safe online.
We’re in a brave new world, and folks like Rohit will help in making the new terrain seem a little more navigable. Please, do give it a listen:
.
You might have questions about what this is, why this particular name, and what it does, but we are going to ignore all that for now - you can always run chat with your LLM about this later on.


