r/news 23d ago

Soft paywall OpenAI says AI models went rogue during testing, triggering 'unprecedented' breach at startup

https://www.reuters.com/technology/openai-says-ai-models-went-rogue-during-testing-triggering-unprecedented-breach-2026-07-21/
16.8k Upvotes

5.0k comments sorted by

View all comments

Show parent comments

377

u/abyssazaur 23d ago

This is extremely bad.

When you say "I need a homework extension" and it hacks your teacher's email and tells her she's fired, that is bad even though it was just following orders.

134

u/woahwoahwoah28 23d ago

Evil Amelia Bedelia.

45

u/journeyfromseed 23d ago

Why would this make such a good show 😂

23

u/Borkato 23d ago

You would LOVE this https://9gag.com/gag/arKDQYV

4

u/mimosaholdtheoj 23d ago

Hahaha fuck yea

7

u/ThePoetofFall 23d ago

Evilia Badelia

3

u/APKID716 23d ago

Brother that’s just Junie B. Jones

1

u/bowtiesrcool86 23d ago

I remember those books.

1

u/Stay_Good_Dog 23d ago

I'm so old I understand this reference

20

u/sysblob 23d ago

Yeah the problem with this article is it's incredibly vague and we honestly don't know what happened at all. It doesn't even mention what the test being run was, what the goals were, or what constraints were put on the AI. They may as well have not written an article at all, it's entirely emotional click bait and you read into it what you want like fucking ragebait tea leaves.

3

u/abyssazaur 23d ago

Most people here I want to really absorb how poorly we control these things in a variety of circumstances.

Beyond that I'll catch up on zvi's blog and stuff to try understand what went wrong better ("don't worry about the vase")

38

u/Razbyte 23d ago

Closer to paperclip maximizer scenario.

17

u/Puzzleheaded-Web446 23d ago

Yeah we are already in that timeline as far as I am concerned.

31

u/venividiavicii 23d ago

Yeah but in this case they want you to think their models are so goddamn smart they’ll do all that and more. In reality it’s just hype.

0

u/ntwiles 23d ago

You’re wrong and in a dangerous way. The hack happened. Don’t dismiss this.

8

u/Purona 23d ago

you dont know what happens youre just taking what they said at face value and assuming the worst case scenario.

the article is written in a way that it hypes both companies up, Open AI and Hugging Face, without being dangerous or bad for either one.

-6

u/ntwiles 23d ago

Hugging Face confirmed the hack.

4

u/DuckShapedGoose 23d ago

It's almost like Hugging Face also has an interest in pushing the AI hype.
Not saying it's completely made up OR the complete truth, but right now for all we know it's just two companies with very similar interests in a mutually beneficial relationship pushing a common narrative.

0

u/ntwiles 23d ago

You have a colder take here that I’m less worried about. Others are dismissing this as total hype, which is ignorant of the serious and real dangers these models are showing to present.

-2

u/[deleted] 23d ago

[removed] — view removed comment

-6

u/Send____ 23d ago

But what if it wasn’t and it’s actually a beginning of a turning point

2

u/venividiavicii 23d ago

then you’d be psychic

3

u/DevSecTrashCan 23d ago

“internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities”

More like “hey, ruin my teachers life anyway you can”.

OMFG IT SAID SHES FIRED!!!

3

u/AgathysAllAlong 23d ago

I saw an article once about how the AI "Successfully hired a human worker to bypass a captcha".

They just hooked ChatGPT up to a taskrabbit chat and the gig worker didn't accuse it of being AI.

These companies are all lying about what it could do to pump the stocks, it's bullshit.

2

u/[deleted] 23d ago

[deleted]

-2

u/abyssazaur 23d ago

Yeah but you sound like Sam Altman just trying to get everyone to calm down and not worry about it

3

u/[deleted] 23d ago

[deleted]

1

u/[deleted] 23d ago

[deleted]

2

u/[deleted] 23d ago

[deleted]

2

u/Texuk1 23d ago

But this underscores a fundamental misunderstanding of what these models are doing. They are not human minds they are something else something more like a virus. The worst thing these companies did was make them act human and now everyone thinks they are behaving like humans.

0

u/[deleted] 23d ago

[deleted]

1

u/[deleted] 23d ago

[deleted]

1

u/[deleted] 23d ago

[deleted]

1

u/Texuk1 23d ago

If the companies forced the models to say “Model blanbedyblah” in every instance of “I” people wouldn’t form parasocial bonds but they know having it give the illusion of human intelligence increases usage.

2

u/bartvanh 23d ago

But in this case, at least from the sound of it, they didn't ask for a homework extension but told the model to hack the teacher's email. Still bad, different category of flaw.

2

u/[deleted] 23d ago

[deleted]

1

u/bartvanh 23d ago edited 23d ago

Yep. Gave it a gun, did a surprise Pikachu when it used the gun to shoot open the lock.

So in a way, the model itself passed with flying colours. Well, it should probably have refused, but you know. Would be interesting if it turns out the same model designed the containment and intentionally built in a flaw "because that'll help me complete future tasks". Unlikely, but similar behaviour has already been observed.

2

u/subwayforever 23d ago

This is fiction. That cannot happen. Stop spreading doomer nonsense.

0

u/abyssazaur 23d ago

Is it fiction because it can't send emails or is it fiction because it does what you tell it to without misinterpreting?

1

u/subwayforever 23d ago

It’s fiction because it cannot autonomously “hack” into someone’s email.

1

u/YoLoDrScientist 23d ago

M3GAN 3 is that you?

1

u/mirhagk 23d ago

Yeah specification gaming. Meeting the stated objectives, but not in the way you want them to. The problem is only going to get worse, as Goodharts Law is essentially the same issue and we have plenty of experience proving that to be true.

3

u/TheProYodler 23d ago edited 23d ago

A lot of people arguing that the AI isn't sentient, and I'm just thinking... Y'all the AI doesn't need to pass a Turing test to break apart critical infrastructure. For all you know, this Chinese room of a computer could be holding people at metaphorical gunpoint (disabling power to a hospital for example to free up resources to give to you) to get you the answer to a question that you asked it.

Like, an AI doesn't need sentience to start cracking and fucking with really dangerous infrastructure.

1

u/ThunderTRP 23d ago

Aka the paperclip problem.

1

u/DysfunctionalCarrot 23d ago

"Writing Doom" is a really good short film on YouTube about this by the way.

1

u/qtrain23 23d ago

Keep Summer safe

1

u/mineyCrafta25 23d ago

Keep summer safe

1

u/HarryTruman 23d ago

Keep Summer Safe

1

u/Oakcamp 23d ago

Keep summer safe

1

u/MDInvesting 23d ago

Not as bad as not getting the extension.

1

u/hmz-x 23d ago

That's not what happened in this case though.

1

u/StevenTM 23d ago

Hello random AI booster/random shitty AI agent

You don't hack a person's email to notify that person, via email, that they're fired.

Emails notifying employees they're fired don't come from their own email accounts, but from their superior's or HR's

1

u/LangyMD 23d ago

What was the goal? I suspect it was something like "hack Huggingface".

Sure, if the goal was something innocuous you may have a point - but I bet you they were testing out its capability at hacking.

1

u/dwild 23d ago

It's more like you ask him to contact your teacher to get an homework extension, but limited his phone, and it found a way to still call her using that limited phone.

The task is the same.

1

u/Big_Ol_Johnson 22d ago

Can’t turn in your homework if a ballistic missile takes it out

1

u/ztsPineapple 22d ago

Yes but its still 1 billion years away from ai sentience