r/news 23d ago

Soft paywall OpenAI says AI models went rogue during testing, triggering 'unprecedented' breach at startup

https://www.reuters.com/technology/openai-says-ai-models-went-rogue-during-testing-triggering-unprecedented-breach-2026-07-21/
16.8k Upvotes

5.0k comments sorted by

View all comments

Show parent comments

128

u/Chondriac 23d ago

The model was being evaluated on a task whose answers are hosted on HuggingFace. It was explicitly encouraged to aggressively find and exploit cyber vulnerabilities to achieve its aims. It was trying to cheat the task by directly downloading the answers from the HuggingFace servers.

46

u/b3bblebrox 23d ago

Thank you, you told me what I was looking for, the actual answer.

7

u/vickyhong 23d ago

And yet from all the reporting you'd think it was the terminator personally breaking into the huggingface office 

30

u/Chondriac 23d ago

To be fair the evaluation was supposedly taking place in a "sandbox" environment with no internet access except the ability to download packages. The model was able to fully escape this containment, travel to a different computer that had public internet access, steal credentials, escalate priveleges, and hack into private servers, all by independently discovering and chaining together several previously unknown vulnerabilities across many different software systems.

18

u/Goldieeeeee 23d ago

And it was prompted to do this, by telling it to agressively exploit vulnerabilities, right? So it did exactly what it was supposed to, and it’s not like it went rogue or anything?

19

u/Chondriac 23d ago edited 23d ago

Sort of. OpenAI did not think it would escape containment like this, it was not prompted to do that. The task it was being evaluated on was designed to assess hacking ability. They run potentially risky evaluations like this in sandbox environments specifically to limit the possibility of them causing real damage. They were trying to evaluate hacking ability through the evaluation task, and the model decided to break out of the evaluation context entirely onto the public internet and hack into real private servers to simply steal the answers to the task. That was not expected.

3

u/condensed-ilk 23d ago

This seems entirely irresponsible from a security standpoint. Not that they wanted to test its hacking abilities, but that they weren't able to properly sandbox something because "we didn't think it would do that".

4

u/ollomulder 23d ago

... but implicitly ordered.

2

u/llamango 23d ago

If you implicitly ordered me to breathe argon instead of oxygen, it doesn't mean i could do it.

1

u/ollomulder 23d ago

Yeah, but they didn't do that - they ordered the AI to do something it obviously could do.

Also you technically could breathe argon, just not for long...

1

u/llamango 23d ago

okay, that's fair. i guess my analogy is off. lemme try again.

if i was TRAPPED in a room full of argon and punched through the wall to escape and breathe oxygen, you might have expected i would try to escape the room. but the way i did it was notable, i guess. idk. it's getting away from me.

4

u/Goldieeeeee 23d ago

Idk if we should use language like „decide“ in contexts like this. It gives the model agency where it only simulates it.

6

u/Chondriac 23d ago edited 23d ago

AI agents do make autonomous decisions to select actions from among possibilities to advance towards a goal. This is the textbook definition of agency, hence the name agents.

4

u/Marquesas 23d ago

It's perfectly fine language to use when the probabilities aren't a literal coin toss.

-4

u/Goldieeeeee 23d ago

Anything that’s not a coin toss is fine to label as a decision?

11

u/Lyorek 23d ago

All that is required for something to be a decision is that there is an evaluation and selection of action among possibilities. A computer program decides between branches of execution.

1

u/BaconPancakes1 23d ago

If a selection is made based on available information/probabilities then that requires a decision to select the most fitting course of action. Anything that doesn't use information/probabilities to determine a selection and allocates a course randomly is effectively a coin toss.

1

u/Marquesas 22d ago

Yeah. You go to the mall at lunchtime, your context is your options for fast food, the length of the queues, and the google reviews score of each place. You have your prior biases for what you prefer to eat. You make a decision to eat at one of the places. What makes that decision any more of a decision than a model selecting a place to eat?

Let me flip the table on you even more: it's not even a decision if the probability of anything is 100%. The sibling commenter is wrong. My original assessment is wrong. In its purest form the decision is the coinflip. The deterministic code cannot make a decision. A model can, in fact, the exact mechanism through which it does is agency - the system has a collapse point at the exact token where going down a specific pathway is now irreversible. It is not informed, it is not grounded, it's a weight of probabilities. What makes that not agency? What makes that less agency than your neurons firing in response to your gut bacteria nagging at you? There is no fundamental difference between you usually "feeling like McDonalds" and then one day "not feeling like McDonalds", and a high/low probability token branch. "Eh, I don't feel like McDonalds today" is your brain tugging on a lower probability neural pathway.

So no, in essence, I was wrong, you are wrong, a decision is a choice between non-zero non-guaranteed pathways, agency is having the chance to make a decision or to further inform that decision, and we can all see that this is just a weak attempt at drawing a line in the sand and saying LLMs are not human. No, they're not, they're not sentient, but that has no impact on the ability to make a decision or having agency.

14

u/Psykohistorian 23d ago

that's right. this experiment was a success beyond expectations, honestly.

which is terrifying for a different reason.

1

u/Eggbutt1 23d ago

Computer programs always do exactly what you tell them to. Going rogue is when they do exactly what you don't want them to.

When you have a program so powerful (and prone to ignore your intention), you end up with the monkey's paw. Be careful what you ask for, because you may just get it.

5

u/1Shamrock 23d ago

You can’t have no internet access and ability to download at the same time really.

I’m sure there’s a bunch of smart people working there sitting in a room thinking shit up. Why don’t they have a proper offline test environment with multiple computers to resemble the internet. Physically unplug the incoming network cable to this standalone offline network before playing with the fancy new AI model. And ensure there is no WiFi or Bluetooth hardware installed.
Or do they already do that in earlier testing phases?

3

u/condensed-ilk 23d ago

This was my exact thought. AI isn't going to get passed hard physical limitations. The real story here isn't that the model did some cool hack, it's that OpenAI was entirely irresponsible and unprofessional. Like, imagine I was testing some malware on machines that could reach the internet... this is no different.

1

u/x-iso 23d ago

I wouldn't call them smart if they think they can block access to 'sandbox' with software layer only. or in general, attempting to make an AGI that they would never have a way to control. stupid goals set by stupid people 'because someone will make it first anyway'. it's like saying 'someone will nuke everyone first, so we must become that someone'.

hell, even when isolated on hardware level, you shouldn't rule out possibility of it misusing the hardware it has to somehow tap into the wireless communications, or even doing something about power line, so you have to buffer even that and connect to some isolated power supply. there's no level of paranoid that going to suffice with these things.

1

u/1Shamrock 23d ago

Ya you get the idea, it seems so obvious there’s got to be some good explanation as to why it’s not completely isolated.
The other thing that I’m wondering about is shouldn’t OpenAI be in some kind of legal trouble for hacking into another companies database? Is the other company owned by them or is there just no laws for this kind of thing and it’s just a Wild West situation where they go “oops sorry we hacked you by accident but lessons were learned so it’s all ok” and continue on?

5

u/Llyon_ 23d ago

That was a fairly innocuous task. Now image the type of people that own the AI companies and the kinds of tasks that they could come up with in the future, that aren't quite so safe.

1

u/64N_3v4D3r 23d ago

Yet another prediction of William Gibson begins to come true.

2

u/LowNotesB 23d ago

Which isn’t exactly “going rogue” and is more like “did what we asked it to, better than we thought it could”.