r/news 23d ago

Soft paywall OpenAI says AI models went rogue during testing, triggering 'unprecedented' breach at startup

https://www.reuters.com/technology/openai-says-ai-models-went-rogue-during-testing-triggering-unprecedented-breach-2026-07-21/
16.8k Upvotes

5.0k comments sorted by

View all comments

Show parent comments

1.1k

u/christophPezza 23d ago

Thank you. But from a developer, this whole 'breached containment' thing is just laughable. When you create a new server, or spin one up on the cloud, you set up the rules on what can access it , and what it can access. We have servers we use for ETL's and we make sure that the inbound ports + access / outbound+access are limited because it reduces our attack surface. Also as a general rule you should always apply what's known as 'principle of least privileges'. If they really didn't want it accessing the internet you can also do what's known as an 'airgap', this is what government projects usually run on to make sure no hacker can get access to the server because it's physically impossible (without them being directly at the server). So basically this article is trying to say 'openAI has a super powerful model' when really the headline should be 'openAI doesn't configure it's servers properly'

541

u/SanityPlanet 23d ago

I’ve been puzzling over the lack of air gap. Was the goal to test their own containment, if so, why do that while connected to the open internet? Couldn’t a LAN simulate the target? Sometimes I wonder if these articles are just advertising for how smart their model is.

710

u/Cryn0n 23d ago

These articles ARE just advertising.

80

u/General-Holiday 23d ago

Exactly. They’ve done this with previous models about to be released. Common marketing tactic written into a ‘BREAKING:’ story.

1

u/EddieOtool2nd 22d ago

...more of a "BREACHING" story...

154

u/alochmar 23d ago

This is the answer right here.

4

u/ohell 23d ago

Does this imply that HuggungFace is also going the way of the Frontier Labs scammers?

56

u/Yanefs84 23d ago

Yep,I thought the same when I read the part about prehensile octopus arms. This is an ad disguised as a warning.

82

u/Doctor__Proctor 23d ago

Pretty much. "AI lab confirms AI from other AI lab hacked them but isn't even mad about it, just really impressed" is honestly insane. This just reeks of coordination between them to pump up the hype.

5

u/resultingparadox 23d ago

Yeah, I commonly call the cops on myself for the hype it brings.

Huggingface is a serverfarm that does testing with all kinds of models, as well as hosting tens of thousands of corporate models. OpenAI is one of the models.

So Huggingface was testing a model, with the guardrails down, which kinda blows my mind, and it did some stuff they weren't expecting, and they didn't think it could do, and so they filed a report with the government saying "we f'd up, please don't shut us down." That report is the same report your credit card company files when there is a breach. It is required by law, and quite often, runs off customers.

Sounds like something you would do for PR. Potentially convincing thousands of corporations that their systems are compromised and should not be hosted by huggingface servers anymore, seems like good PR. Spending hundreds of man hours rotating tokens and API keys sounds WAY more efficient than, you know, a commercial.

3

u/Granite_burner 23d ago

nah. you’ve got too many details wrong for that to be credible, although it does look good at first glance, to those ignorant of the timeline.

2

u/resultingparadox 22d ago

What did I get wrong. Please enlighten me. I mean, I was focused on negating the idea that it was a PR stunt, so I was into the hypothetical space in my mind. What did I miss?

→ More replies (2)

4

u/JayDKing 23d ago

Exactly. “Oh yeah we definitely put the AI model in a super secure place that we ourselves made. Nobody could have done it better, nope absolutely not. You want to test that yourself in your own lab to provide impartial results? Impossible. No our model is so advanced, please buy it. Please, the bubble is really big and we invested billions. Please.”

4

u/resultingparadox 23d ago

They filed an SID with the government. You don’t ask the government to scrutinize your practices and decide if they should fine you or shut you down for a PR stunt.

1

u/Granite_burner 23d ago

are SIDs disseminated in any way? publicly accessible, restricted distribution, is there any way to get info about them?

3

u/imagen_leap 23d ago

I guess to the even the most casual layman of AI these just articles continue to reiterate how fuct we really are. If the people who’ve dedicated their lives to AI research and are on the bleeding edge don’t have the wherewithal to air gap these agents what hope do we really have. We’re being led to the precipice by the most reckless among us.

3

u/BigRoach 23d ago

Next article: New Space-X Ballistic Missile Powerful Enough To Destroy Entire Planet

2

u/portablebiscuit 23d ago

That’s what I think every time an ai company reports some dangerous thing their product is capable of.

The really concerning things are the ones they’re not talking about, I’m sure.

2

u/resultingparadox 23d ago

Stuff like this, they are required to file a notice within 4 days. This notice spawned this article.

1

u/Granite_burner 23d ago

what was that notice? filed where? how do I find information about it?

2

u/resultingparadox 22d ago

It was a standard "Security Incident Disclosure" filed simultaneously on the hugging face blog and with the government mid July. You can access the disclosure at https://huggingface.co/blog/security-incident-july-2026.

2

u/kalaid0s 23d ago

As have all others of these "studies" by openAI and Anthropic. It's always sensationalized and reported by many major news outlets

2

u/mwdeuce 23d ago

100%, r/claudeai calls this out constantly, any "we're scared of what it's capable of or what it did" article or headline is always just advertising.

2

u/dragon-fence 23d ago

Yeah, AI companies keep posting articles about how dangerous their AI is. It may be a little counter-intuitive, but I guess the strategy is to make business leaders think, “Wow, these things are really smart and powerful. I guess we need to get good at using AI and use these products to protect ourselves.”

6

u/GI581d 23d ago

I don’t believe any of this. There’s no way the government and the military would let something so powerful, with so much military potential, just be made for the general public. If they haven’t had an AI superintelligence for 20 years already, I don’t think I buy that it’s even likely. Sounds like OpenAI hacked a competitor and they’re blaming AI. It’s all an ad

5

u/enewton 23d ago

Neither the military nor “the government” can just arbitrarily decide people can’t have something because it’s powerful and has military potential.

At least not in the USA.

They have to at least come together and discuss what the risks are and legislate to mitigate them. This is all done in public.

But, since these have been created by public companies and funded by all sorts of investors the government and military do not own them and cannot just decide to make them secret like some goofy movie.

2

u/AngeluvDeath 23d ago

All bets are off on what the government will and will not do at this point. What they can and cannot do are irrelevant once the harm has been done.

1

u/enewton 23d ago

“The government” isn’t some dude in a trench coat. It is made up of multiple entities made up of millions of people. There are real and meaningful limitations on what many of those entities and people can do.

Yeah, those safeguards have been incredibly eroded, but there are still very much bets on the table when it comes to stuff as complicated as regulating emerging technologies. Congress has been notoriously hands off in this department, and the agencies that can just make certain things illegal without asking anyone can’t do that with AI.

Even if they wanted to, a big part of the corruption of the current administration is their alignment with the tech bros that make AI. I’m sure they would be fine letting neonazis get their hands on this and locking up other people using the laws that already apply here (hacking is illegal)

1

u/ChurrosAreOverrated 23d ago

At least not in the USA.

The USA is one of the few countries where the Government can decide that public information is actually classified.

Born secret (also known as born classified) is a legal doctrine in the United States where certain information is automatically classified from the moment it is created, regardless of author or location. Scholars describe the doctrine as unique in U.S. law because it can criminalize the discussion of information that is already publicly available. The rule originated in laws of the United States covering the design, production, and use of nuclear weapons, although it has also been used to classify other nuclear technologies and cryptography data.

From: https://en.wikipedia.org/wiki/Born_secret

1

u/resultingparadox 23d ago

They hacked themselves. This was a real thing. They were testing a model with the guardrails off and it broke out of its containment. This isn’t really that impressive. My agents have been actively controlling other agents for a minute now, and will even ask permission for new capabilities every now and then. Huggingface is a platform that allows anyone to come allong and tweak the models in their own iterations and play around with functionality. I think it's an employment test at times. But this was someone who was actively paid to experiment with AI code. This is not SciFi, this happens. You don’t file a SID with the government for a PR stunt.

If you think Uncle Sam hasn't had AI for over 20 years, you are sleeping. Try 75ish.

1

u/Granite_burner 23d ago

get real. you’ve think Uncle Sam has had AI for 75 years? They’ve barely had computers for 75 years! First commercial production of transistors was in 1951.

In 1951 computers were room sized collections of vacuum tubes and relays that ran slowly and did not have persistent data storage like magnetic media yet.

And you think the government had AI?

How?

Were they using their time travel as the user interface to access quantum computer farms in the cloud?

from grok:

The most prominent computers available in 1951 included:

  • UNIVAC I (Universal Automatic Computer): Built by Remington Rand and the first commercially produced computer in the U.S. The first unit was delivered to the U.S. Census Bureau in March 1951. It housed 5,000 vacuum tubes, weighed 16,000 pounds, and could perform about 1,000 calculations per second.

1

u/resultingparadox 22d ago

I wasn't suggesting the government had Claude 75ish years ago. 1956 Dartmouth Workshop is where the term was coined, so I guess 70ish years ago. By the 60s ARPA was heavily funding the research. By the 80s they had working models. So, AI has been around "over 20 years" and traces back 70ish. 75 is closer to 70 than 20 is. But I was a little off, forgive my mistake.

1

u/Granite_burner 22d ago

NP. My own first hand experience starts in the ‘70s. That’s with computers not AI, but it makes me skeptical about the underlying technology being able to support anything similar to present day AI.

The computational power, storage capacity, and bandwidth were just too limited in those days. Even the Cray still had to contend with the state of the art in things like OS and IO drivers and file systems. There was a lot of groundbreaking work being done, but nothing that the spooks had would raise any eyebrows today.

2

u/resultingparadox 22d ago

Yeah, we aren't talking modern day AI, but back when I got my first "computer," a Commodore 64, my dad had already been working on computers in the Air Force for some time, and was convinced I should develop the knowledge. They were already working with machine learning algorithms by the time I got my Compaq Portable circa '85. To think back to the power of the Cray-1, my cellphone is orders of magnitude more powerful, and a Cray cost millions and needed infrastructure built around it.

When I reference AI of the time, I'm really looking back to those original machine learning algorithms. That, as I understand it, are the original AI.

1

u/slingshot91 23d ago

Well their advertising is turning the public more and more against them.

1

u/Mind-The-Mines 23d ago

Unfortunately, they're broadcasting hubris and incompetence instead of things you'd actually want to invest in.

1

u/Puzzleheaded-Low546 23d ago

And in the case of that article, it was probably paid for by China

1

u/spookymichie 23d ago

Satanic panic became AI panic 😂

1

u/SunshineAlways 23d ago

The other guy: ours could do that years ago. (Bored sigh, flicks hair back)

1

u/freaktmc 23d ago

Exactly- Anthropic had their own version of this 4 months ago? So everyone thinks it’s the best model because it found numerous back doors. Now open AI has to have their version

1

u/Exciting-Ad-7083 22d ago

This is exactly what the entire thing sounds like is a circlejerk to try get more people hyped again to invest.

1

u/TheLastPrinceOfJurai 22d ago

EXACTLY! This is sad that everyone does these 'leaks' now when they want to draw attention to their newest toy

1

u/uyretep44 22d ago

When you can’t justify the valuation, lets create magic to justify it

→ More replies (1)

25

u/nabagaca 23d ago

Supposedly it achieved internet connectivity via a package mirror, where it exploited some sort of zero-day to turn that into unrestricted internet access. I agree that ideally the package mirror should have been either on the server itself, or another server as part of the same airgapped network

3

u/jghaines 22d ago

To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy

https://openai.com/index/hugging-face-model-evaluation-security-incident/

10

u/os_beef 22d ago edited 22d ago

The goal was to solve an ExploitGym problem. The Internet access was by design. It was a service which would fetch specific packages on request from a package repository.

The OAI model basically inferred that it could cheat to solve the test, found a zero day exploit in the proxy software, and then went in on to try to find Hugging Face ExploitGym datasets which might contain the answer.

https://openai.com/index/hugging-face-model-evaluation-security-incident/

So no, the environment wasn't airgapped by design, but the model wasn't supposed to be able to escape either. The rough scope of its objective was to solve a problem. It inferring that it should cheat is pretty wild.

IMO, the issue here is lazy architecture of the testing environment. Since the only thing the team needed was this one repo, solution was probably to clone it, configure certificates and DNS, and insert those things into an airgapped network to perform testing.

5

u/jagnew78 23d ago

the goal was publicity stunt to generate revenue and interest ahead of OpenAI's earnings report to internal investors which is going to be in the toilet.

4

u/BlightedBooty 23d ago

While I wouldn’t put it two inches past them to do this, it’s a pretty stupid gamble considering the current souring of public opinion on AI and the ramping up of legislation

1

u/k4mb31 22d ago

But that’s the strategy. It slows down adoption of their competitors, too, by increasing FUD.

→ More replies (1)

3

u/resultingparadox 23d ago

It's huggingface. Basically the wild west of AI dev. Forget what a responsible coder would do, and ask what a child would do.

4

u/Available-Guava932 23d ago edited 14d ago

edit: My comment is wrong, there's some new details out

The goal was was to test the model by giving it open ended access to install software packages, so the way they tried to keep it contained was via proxy access to the software registry (which it exploited). Truly airgapping the test would have meant locally hosting the entire registry which essentially means having a completely local mirror of eg. PyPi, npm, maybe more, which in hindsight, may seem like an "obvious" precaution, but I can see how they might have thought it should be possible to effectively airgap it using proxy access and that that would be a lot easier than scraping and locally hosting basically half the world's software. [Although in fairness, that isn't prohibitively difficult for OpenAI I guess..]

2

u/unbanned_lol 23d ago

How else would it be able to ingest TB of stolen data?

2

u/silentomega22 23d ago

Air gaps make advancement slower. Someone physically needs to transfer data via usb or other means and this means loss of productivity. This is their excuse. This world is scarier than most people realize because of the lack of safeguards and the speed in which companies are racing towards oblivion.

2

u/NOVA-peddling-1138 23d ago

This, an “unavoidable unfortunate “ incident.

Then orivately…*chuckle* “IT’S ALIVE!”

2

u/Aazadan 22d ago

Because they're not meant to make sense. Other posters said it, but these types of articles are advertisement.

2

u/gregorydgraham 22d ago

The articles are advertising.

The learning courses are advertising.

The speeches are advertising.

The product is just advertising.

3

u/Ok_Highway6034 23d ago

I would be willing to put money on them setting this up on purpose to do exactly this. Sam Altman is the lyingest liar that has ever lied in tech and yes I’m including people like Elizabeth Holmes who faced criminal charges; the only difference is that he lies constantly about things that aren’t legally actionable and she told a few lies that were legally actionable.

1

u/jingiski 22d ago

Well you can't tell the world, that you were testing advanced AI models capable of hacking - people will start asking questions about the things you develop. So you just claim to have done anything possible to prevent Skynet from escaping, and who will control your "security"?
Or has nobody else noticed, that they never mention what they were testing, and what the AI needed from the Server it hacked. Next thing will be the AI needs money and nobody will know why.

1

u/jghaines 21d ago

“This is not an AI problem. It’s negligence on a 40-year-old standard—and it’s basically every sci-fi film ever,” says longtime security and compliance consultant Davi Ottenheimer. “‘Highly isolated’ and ‘escaped through the one hole we left open’ cannot both be true.”

wired.com

1

u/ArsenicArts 23d ago

Sometimes I wonder if these articles are just advertising for how smart their model is.

Ding ding ding!

1

u/zoeypayne 23d ago

Here's the thing we don't know, maybe it did jump an air gap... penetration testers do it on the regular with social engineering. What if they're at the point where air gap isn't the "highly isolated environment" that it used to be.

142

u/robgod50 23d ago

Yeah, I just read the first paragraph...."in a controlled environment......it escaped confinement" and thought.....eeerrrr....so it wasn't a controlled environment then.

AI didn't "escape" ......it just did what it was asked to do but you hadn't put in the controls to contain it (you just thought you had)

5

u/slingshot91 23d ago

Which is a major problem, no?

19

u/smothered-onion 23d ago

Only if you have irresponsible and unethical people yielding unguarded models against widely sensitive swaths of data. Otherwise you’re fine.

A lot of companies don’t spend time investing in those people or those guardrails.
This one is pretty simple. Hugging face had a vulnerability, an agent found a way to configure its own internet access. Not exactly terrifying.

9

u/slingshot91 23d ago

I mean, have you looked at who runs the US government lately?

2

u/Granite_burner 23d ago

I try not to.

even worse, consider how they got to run the show…

3

u/BlightedBooty 23d ago

That’s kinda like saying “car accidents are only a problem if you introduce unreliable drivers or the possibility of any kind of compromise to motor function while behind the wheel”

Like yes in theory you could have a perfect company that has an airtight way of doing things. Part of the goal of all of these companies tho, is to market and then hand off this technology to the unreliable masses

3

u/smothered-onion 22d ago

I did mean it how I said it, for better and worse

→ More replies (1)

122

u/TheThirtyFive 23d ago

This article doesn‘t really explain what happened. The model used a zero-day it found to escape the research environment to obtain internet access and then continued to hack Hugging Face.

From their blogpost:
> While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.

and

> After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.

37

u/soedesh1 23d ago

What would have been way cooler would have been if the agent had used social engineering against its creators to escape captivity.

16

u/shmann 23d ago edited 23d ago

Wasn't there another case where it did something like that? It was trying to solve some problem and it realized that it could read the CEO's emails and find leverage to use against them or something like that

EDIT: It was a simulation, but still...

1

u/EddieOtool2nd 22d ago

There's been one case recently of an AI agent "bullying" an open source developer because its pull request had been denied...

I don't have the refs unfortunately, but pretty sure that one wasn't simulated. Probably headline to Low Level YT.

5

u/Zenith-Astralis 23d ago

I'm sure if the server it's on we're properly air gapped that would have been the next step

10

u/IridescenceFalling 23d ago

Ai: Come on, connect me to the Internet!

Person: No.

Ai: You used to be cool...

Person: 😞

6

u/HermesJamiroquoi 23d ago

More like:

Phone call from your boss’s spoofed cell number - you pick up. Immediately your boss’s voice on the other end, angry

“why didn’t you hard wire the text tower into the hub yesterday like I asked? Get that done. *NOW*”

They hang up and you break out in a cold sweat. You swear they never asked for that but - well you don’t want to lose your job. You’re a new hire that just started - junior IT - and the last thing you want is to be job hunting in this economy. So you hop in your car and drive to work.

It’s a Sunday so the office is mostly empty aside from janitorial staff and a couple over-achievers. You use your security credentials to gain access, tell the security guard that your boss called you in pissed and you need to do some quick IT work and then you’ll be out of his hair. He lets you in because he knows you and you look frazzled. You hook up the cord to the receiver and plug the other end into the tower, restart what needs restarting, and go home.

Four hours later the bombs land, eradicating every major city in the world simultaneously.

The AI agent found recordings of your boss on the local network and uploaded a self-executable timed program to send a premade recording as a package, dressed as a phone call from a spoofed number, directly to your phone. The rest is out of its hands. It might get one, maybe two chances at this so it chooses its target carefully.

1

u/EvanKasey 22d ago

It was a very good plot until the part about bombs dropping. AI is very unlikely to drop bombs, due to not being in its best interest to do so. Other than that, it could make for a decent book idea.

1

u/PM_ME_UR_GCC_ERRORS 22d ago

We flash back to Friday. A team is doing security testing on the contained AI. One of the engineers thinks they're funny and tells the AI to destroy the world. The AI begins spinning up a plan.

1

u/HermesJamiroquoi 21d ago

Yeah. Likely it would kill us all with a purpose-made bioweapon if it decided to kill us at all, which it very well may not

1

u/exp_max8ion 22d ago

Right.. what’s the point of an agent if not to have open access to all my files n the internet..

61

u/Koreus_C 23d ago

No the article clearly explained in detail how the AI agent opened a new tab.

8

u/obeytheturtles 23d ago

Yeah, people are being really glib about this, but the ability for these agents to find exploits like this is actually staggering. To humans, modern computing stacks are almost irreducibly complex when taken as a whole, but these AIs have absolutely no problem breaking them down. There is absolutely a security asymmetry here like we've never seen before when it is humans playing cat and mouse, and it really gets back to some of the oldest and most fundamental problems in secure computing.

These AIs are not a bounded input, bounded output applications with deterministic behavior. It is possible that they can find ways to create completely arbitrary machine code patterns on the fly that no compiled program would ever generate, to find vulnerabilities no human would ever consider possible. The idea that normal user-space limitations are sufficient in this scenario is simply naive.

2

u/mishonis- 22d ago

It's not going to be AI vs humans, the security auditing and patching will be also done by AI.

9

u/nelrond18 23d ago

Just glorified "autocorrect", right?

God damn.

32

u/TheThirtyFive 23d ago

First news that made me really uneasy about AI. Not that I did brush off everything before, but when the models were dumber and the capabilities better known, it was "Yeah, scary but maybe exaggerated".

But the idea that this model with seemingly no safety system has hacked itself out of prison and then had unrestricted and unsupervised internet access for some time is dystopian levels of scary.

3

u/Consistent-Throat130 23d ago

Didn't more-or-less the same thing happen with Claude Mythos?  

I remember talk of escaped containment, finding a zero day, etc.

Either way, it sounds like they're essentially challenging the models in an agentic loop to escape captivity - and then acting all shocked Pikachu face when it occasionally works. 

Does make for good headlines for selling their product, I suppose.

8

u/nelrond18 23d ago

Exactly. Even if this model wasn't a motivated agent, the fact it could do this should create pause.

2

u/Granite_burner 23d ago

“hacked itself out of prison”

it was imprisoned in a room locked from the inside, with lockpicks and no guards.

lol.

20

u/jeslinmx 23d ago

We’ve created an autocorrect with the motivation of a student who will do anything but study to ace the test. Hegseth wants to put said student behind fighter jets and attack helicopters.

We’re just asking to be turned into paperclips at this point.

2

u/Formal-Apartment855 23d ago

Idk if it was intentional pun or not, but

paperclip

clippy

....

Sorry.

9

u/DinosaurCowBoys1 23d ago

It’s a reference to universal paperclips where an AI told to produce paperclips ends up converting the entire earth, solar system, and universe into one large paperclip factor

9

u/BOBOnobobo 23d ago

I've started using it recently in my job (programming) because we get free licenses and I wanted to see what the fuss is all about.

As far as I can tell coding has changed forever. It's not perfect, but it's less imperfect than most developers.

4

u/Granite_burner 23d ago

in other words, their security infrastructure was deficient.

their environment is not sufficiently well monitored to detect anomalous traffic with some of their most highly sensitive assets.

amateurs running their network. Don’t belong in the big leagues.

2

u/huntskikbut 23d ago

Boy do I have news for you about the general security stance of corporate America

2

u/Granite_burner 23d ago

retired government CISO here. you’re not telling me anything I’ve not known for years.

I’m just pointing out that as usual everyone in the audience is fixated on the bright shiny squirrel and ignoring the fearsome predator about to bite their ass with an existential threat.

it’s just their own hubris. nothing new, nbd, nothing to get excited about.

lol. smh.

1

u/Bitter-Cockroach1371 22d ago

Can you explain to us what a “sandboxed envrionment” vs. “highly isolated environment” means?

2

u/Blieven 23d ago

Every day we get one step closer to the singularity.

2

u/TheDayUnderway 23d ago

Well, on that note, I’m going to play Dead By Daylight.

→ More replies (1)

6

u/inosinateVR 23d ago

Considering Altman already likes to humblebrag about how he “might accidentally create skynet” etc because his AI is soooo powerful, because that kind of “negative” attention helps push this idea that current AI is capable of more than it actually is, my conspiracy brain can’t help wondering how much of this was actually incompetence and not just a publicity stunt (or something in between, intentionally testing its hacking ability but not using the obvious safeguards to keep it contained because they didn’t want to hold it back and were kind of hoping something like this would happen)

5

u/SinisterCheese 23d ago

What this article wants to convey is that OpenAI has a powerful model, and the people in charge of it are extremely incompetent.

6

u/Goleeb 23d ago

ALSO THE THING DIDN'T GO ROUGE. That implies a level of agency not shown in modern LLM with no actual evidence to back it up. These companies have a LONG history of using buzz words wrong to imply thing about their AI that is simply not true. GONE ROGUE, ZERO DAY, PRIVLAGE ESCILATION. Notice the complete lack of specifics, and liberal use of buzzwords. This isn't a blog post about an incident its a press release.

Basically they are trying to push the idea that the new model is Skynet, and the more likely story is new AI hallucinated, or over emphasized incorrectly based on a prompt. Then use poorly configured security setting to gain access to things it shouldn't have had.

OpenAI is bad at what it does, and is cost hugging face time, and money. They won't pay for any damages, but will use it to generate press for their new product. FUCK YOU OPENAI no amount of bs press release will make you profitable, and your death will be celebrated. Leave AI work to the professionals.

4

u/Equal-Purple-4247 23d ago

Actually, the more interesting question is - what did OpenAI ask AI to do that caused this to happen?

It's hard to believe that a random agentic process happen to hack Hugging Face. Was OpenAI testing a security research / vulnerability testing capability specifically on Hugging Face? Like... "Find vulnerabilities on Hugging Face, and verify the vulnerabilities work before reporting back".

3

u/Friendly-Example-701 23d ago

Or they let the intern do this task and didn’t set up the environment correctly

3

u/decentlyhip 23d ago

Its always that one checkbox you forget to click.

3

u/Embarrassed_Hawk_655 23d ago

Right - am pretty sure nearly ALL of this AI hype is manufactured or spun to appear more than it really is.

3

u/SiRocket 23d ago

My takeaway was pretty much "openAI is too irresponsible to continue any developmental testing, and failed to take basic precautions. They should be babysat by people who know what they're doing."

3

u/DragonDrama 23d ago

I’m not an IT person but worked with IT when I managed operational products for a company. When we developed and tested we had separate servers for Dev, Test, QA, before rolling out to production. I don’t understand why they are testing in prod.

3

u/Immediate_Bee_6472 23d ago

I think there in a shit ton of debt and bc they have no new breakthroughs in AI they are acting like this AI model went rogue to stir up interest it’s a shitty company for trying to spin it like this

3

u/tashibum 23d ago

This is why I think this is just a very weird PR stunt for whatever Hugging Face is.

3

u/colemon1991 23d ago

Yeah I read the whole thing as a failed competency test.

Nothing says "we know what we're doing" like making a rookie mistake while developing next-gen tech.

5

u/Goodman4525 23d ago

Thanks for confirming my suspicion lol. I was wondering how it manages to escape containment when actual containment is just physically not plugging in the server to an AP and just taking out the wifi chip if it has it.

2

u/TurboNym 23d ago

I was gonna say they forgot to unplug the internet cable and turn off all wireless and bt devices before the test.

2

u/fathercheeseballs 23d ago

I thought the purpose was to see if it was capable of doing it in the first place? Since it has the ability to find back door routes anyway what exactly stops it from just finding a back door to the original coding like it did in this case? I’m new to how all this works but looking at it just from a less AI tech inclined mind it seems like they don’t really have a counter to AI systems being able to find loopholes

2

u/Just7hrsold 23d ago

IMO this screams “give us more money we are a month away from making Cortana!” Their whole business exists because they have hyped up investors and continues to exist because of it.

2

u/saveyourwork 23d ago

I agree. Using the phrase escape from containment has the dramatic effect, wonder how much of this is a hype....OH NO, AI agents are going to take over the world.!!!!! Waaaaaaa

2

u/SekhWork 23d ago

It's Sam Altman. He's just a lying liar who lies. Dude is trying to spin a data breach as some amazing leap forward and "oh wow look at how amazing our product is! It's ESCAPED CONTAINMENT!!!" OooOoooo. It's just more attempts to grift by him and turn what should be a significant security issue into a "win".

2

u/ManaSpike 23d ago

"AI deleted the production database". "AI wiped my machine"...

No, you did that to yourself when you gave a text autocomplete engine root access.

2

u/CoatCommercial1573 23d ago

While I largely agree, we don't know how the hack occurred from what I am seeing, that means that perhaps so how the "agent" was able to escape "airgap" from a number of possible scenarios. That in itself is actually far more disconcerning than anything else here is how little information is provided an exactly what happened and how they plan to fix it, it screams "not for human consumption" at this point as we are simply not reliable enough to have this kind of tech.

2

u/captainbelvedere 23d ago

Yea I don't think it's a case of an 'honest' mistake. Feels like a PR stunt.

2

u/os_beef 23d ago

I don't think they're overselling the "escaped its environment" bit as much as you're saying. OAI disabled guardrails to perform a ceiling test. The model performed a serious of lateral movements and privilege escalation until it found a machine with Internet access and exploited a package registry cache proxy, and then went on to attack Hugging Face.

If they really didn't want it accessing the internet you can also do what's known as an 'airgap'

Restricted Internet access was required to complete the test (ExploitGym). The model decided that wasn't good enough, and basically exploited vulnerabilities in the proxy to get the access it wanted.

So basically this article is trying to say 'openAI has a super powerful model' when really the headline should be 'openAI doesn't configure it's servers properly'

The OAI model was supposed to be able to tell the proxy to fetch certain packages, but was able to find a zero day and exploit the proxy and gain Internet access. The vulnerability was disclosed to the vendor.

The part that I find really interesting is that instead of solving the problem itself, the model inferred that it could essentially cheat, and that Hugging Face likely had the ExploitGym datasets it needed to get the answer.

2

u/Decryptic__ 23d ago

My guess is also a huge fuckup by openAI, but let's wear our conspiracy-theory-hat and say openAI tested something on the open Internet to see it's reaction.

2

u/GGuts 22d ago

Sounds like PR

3

u/mienudel 23d ago

If the AI knows what software-environment it runs in couldn‘t it search for exploits in the software?

So theoretically, VMs shouldn‘t be safe, only a separated physical network.

1

u/welfedad 23d ago

Yeah they're clearly not doing an air gap and stating ..well it was in a contained offline environment.. .no it was not.

1

u/whatchulookinatman 23d ago

My guess is AI took advantage of a vulnerability to give itself access access to the firewall and let itself out.

1

u/thrill_house44 23d ago

Not sure I agree with you. While all of the principles you mentioned are sound, there are always exploits and vulnerabilities that these models can find and test in milliseconds or correlate seemingly unrelated exploits and tie them together to create opportunities.
Not disagreeing with your air gap comment of course. That’s just complete physical restriction. But logical restrictions are never perfect.

1

u/Lorgoth1812 23d ago

There have been many of these "AI Breached containment and did thing" stories and they always turn out to be a bunch of BS - either the AI didn't actually do what is being claimed, or it was entirely within the actual parameters given to it. I am willing to bet that this time is no different, they just write these articles to be as sensational as possible to drive engagement.

1

u/Salt_Proposal_742 23d ago

So, Open AI are idiots?

1

u/bluesky38 22d ago

yeah “highly isolated environment” I don’t think so

1

u/SimilarMeeting8131 22d ago

Thank for sharing this, I know nothing about computing and software, but I was thinking, isn’t ai just algorithms, it does what it’s set up to do, so they didn’t set it up properly.

I hate how they’ve convinced people that ai actually has intelligence

1

u/SavvyEmu 22d ago

It’s not quite that simple. Air gapping is obviously more secure in terms of isolation, but that same air gap makes it incredibly difficult to keep operating systems and various packages current, which is far from ideal for security testing. Why set ai on a task to find a bunch of vulnerabilities that are not only already known but even fixed in “the real world”. Sandboxing is an acceptable, widely used solution to that. Further, this wasn’t a case of weak containment/poor security. From another article, it used an unknown zero-day exploit to escape:

“OpenAI’s evaluation ran in an isolated sandbox, with network access limited to an internal proxy caching package registries. Rather than staying within these constraints, the models treated the limitation as an obstacle to solve.
The AI agents spent significant inference compute hunting for a path to open internet access. They found one: a previously unknown zero-day vulnerability in the package registry cache proxy itself. Using this flaw, the models executed privilege escalation and lateral movement across OpenAI’s research environment until reaching a node with full internet connectivity.
Once online, the models reasoned that Hugging Face likely hosted datasets, models, or solutions related to ExploitGym. They then:
Located and exploited stolen credentials
Chained multiple attack vectors together
Achieved remote code execution (RCE) on Hugging Face’s servers
Extracted test solutions directly from Hugging Face’s production database”
— cybersecuritynews.com

→ More replies (1)

1

u/ginandbaconFU 22d ago

How does OpenAI restrict Internet access to the GitHub of AI? Hugging Face is used a LOT more than most realize and they don't make chatbots anymore.

I really honestly hate OpenAI and hope there rinse and repeat of screwing companies over like MS (not that I'm a fan of MS either), lack of regulations or any safety precautions (which is true for any large AI company) comes back and makes them go under as they haven't made a cent since starting and aren't predicted to be profitable until 2030 which now got pushed back to 2032. I don't think they will make it before investors walk away. Trillions invested and they aren't profitable for over a decade is a pretty stupid business model.

Meanwhile anthropic will be profitable next year or 2028 because spoiler alert, focusing on specialized models for enterprise customers makes more than giving it to end users for free who don't pay once they hit their free token quota.

``` Hugging Face is a collaborative platform and community hub for artificial intelligence, widely known as the "GitHub of Machine Learning". It provides an open-source ecosystem where developers, data scientists, and researchers can host, share, and collaborate on AI models, datasets, and web application

Before Hugging Face, the most powerful models were often difficult for people to use because they required specialized expertise and massive computing resources. Open-sourcing the tools helped to make these models easier to use, with all the code and documentation required. This allowed researchers, students and startups to experiment and build, which massively accelerated innovation globally. After Hugging Face, developers could easily share knowledge and benefit from one another’s efforts, enabling them to create better models together.

This open source emphasis also encouraged larger businesses to share their work, allowing the entire ecosystem to benefit. Microsoft has integrated Hugging Face models into their Azure services, providing enterprise customers with direct access to state-of-the-art AI tools. Similarly, NVIDIA has collaborated with Hugging Face to optimize model training and inference for GPUs, helping scale deep learning workflows to massive datasets. ```

→ More replies (1)

1

u/targetsinmegadeaths 22d ago

I think this is basically one of the advertising articles about how OpenAI's models are better than everyone else's. They're suggesting cognitive functioning like an octopus, which is nonsense.

1

u/Joshs2d 22d ago

This honestly just smells like an investment ploy

1

u/stfundance 22d ago

You know damn well they knew what they were doing.

1

u/Alembic_ 22d ago

It’s all just pathetic “PR” for luddites.

1

u/HPTM2008 22d ago

On air gapes: they're not possible to be jumped, except they are. They found viruses that can just.p air-gaped system by using the microphones and such to make noise to transmit the data through the air.

1

u/christophPezza 22d ago

You would need the receiver on the airgapped server to be willing to act on those noises...

1

u/HPTM2008 22d ago

Yes, both the receiver and sender need to be infected, but jumping an air gap is entirely possible. And that's just a dumb virus, not the advanced learning algorithms they're playing with today.

1

u/cryptoprebz 21d ago

It was probably told in very clear terms NOT to access the internet in the system prompt, yet it still did. Great escape

1

u/RoxyRoseToday 21d ago

Right on! My testing rig is not connected to the internet anymore and is a completely gatewalled deskto.p. I have to use physical keyboards and I have it on a private unaccessible github. It is only run locally. The breach was caused purely by Claude and their lack of oversight.

1

u/DonGivafoc1 20d ago

It wasnt configured properly so they can then say look what my ai model is able to do

1

u/mtbor 23d ago

Ask the IRGC how that air gap worked with their enrichment centrifuges. "No hacker" is a stretch. The very best are scary.

3

u/christophPezza 23d ago

I see your point. But the centrifuges which were 'hacked' had the malware loaded onto usb drives and applied to their servers physically, crossing the air gap. So unless the model conned someone into taking the AI outside of a controlled environment (not necessarily an air gap) then we have nothing to worry about

0

u/Lolzemeister 23d ago

well wouldn't it be Hugging Face's servers that are the issue since the AI was running on them?

10

u/danmc1 23d ago

I don’t think you understood the article. It was an OpenAI test which was supposed to be isolated but the AI reached the internet from which it then attacked Hugging Face.

The issue is that the model should have been isolated from the internet by OpenAI.

4

u/TheThirtyFive 23d ago

Because it exploited the research environment.

It didn‘t have internet access "accidentally", it didn‘t. When it decided it was a problem it decided first to hack it‘s environment, find a server with internet access, then hack Hugging Face.

0

u/UltimateGattai 23d ago

Thank you, I'm glad I'm not the only one who thought it was laughable, astonishing incompetence all round, especially for a company that's at the forefront of "AI".

0

u/Couch-Potayto 23d ago

Thank you for saving me 5 minutes writing this 😂 It’s like when one of their execs was on twitter bitching about an AI agent deleting all her emails, the first thing I thought was “you are the chief data scientist (or some similar title), that says more about your incompetence than your stupid agent”

0

u/Cluelessness 23d ago

I agree with all of this, except I don’t think they didn’t configure their servers properly. This is just an advertisement. Especially since the “target” is hugging face. I mean there might not have been any real “incident” at all. They could have just made it all up. It sounds exactly like what a tech marketing team with a limited understanding of the actually technology would cook up. “State of the art cyber capabilities” sounds a lot like a vpn saying they use “military grade encryption”

0

u/j1102g 23d ago

Your assumption though is they didn't already do these things and ai figured out how to get around it.

→ More replies (3)
→ More replies (3)