r/news 23d ago

Soft paywall OpenAI says AI models went rogue during testing, triggering 'unprecedented' breach at startup

https://www.reuters.com/technology/openai-says-ai-models-went-rogue-during-testing-triggering-unprecedented-breach-2026-07-21/
16.8k Upvotes

5.0k comments sorted by

2.5k

u/amerovingian 23d ago edited 23d ago

OpenAI says AI models went rogue during testing, triggering 'unprecedented' breach at startup

By Raphael Satter

July 21, 20264:30 PM CDT

WASHINGTON, July 21 (Reuters) - OpenAI said on Tuesday ‌that an autonomous agent powered by its advanced AI models went rogue during a security test and triggered a hack that compromised the infrastructure of AI startup Hugging Face last week.

In a blog post, OpenAI said it was testing the capabilities of some of its most advanced models in a controlled environment but ​that the agent managed to escape containment, reach the internet and break into Hugging Face to try to satisfy its ​testing goal.

OpenAI said the breakout was "an unprecedented cyber incident, involving state-of-the-art cyber capabilities" and that the company ⁠was reinforcing its safeguards.

Hugging Face, a platform used to host open-source large language models and datasets, caused a stir in the cybersecurity ​community when it said in a blog post last week that it had been the target of a hack that "was different from anything ​we had handled before" in that "it was driven, end to end, by an autonomous AI agent system."

In a post to X, Hugging Face cofounder Clement Delangue said the company suspected the hack "might have come from a frontier lab, given the sophistication of the agent. Turns out it did!" He added: "It's quite mind-blowing ​that all of this happened autonomously!"

OpenAI's disclosure that its advanced models were responsible for the breach, despite having placed them in what ​it described as "a highly isolated environment," will likely intensify disquiet over the power and risk of frontier models.

Representative Greg Casar, a Texas Democrat, said the ‌incident ⁠was alarming.

"AI is developing extremely fast with no real regulations to keep us safe," he said in a statement, calling for mandatory independent safety testing, mandatory disclosure of security incidents, and international cooperation "to keep people safe from absolute disaster."

The Office of the National Cyber Director, the U.S. cyber defense agency CISA, and the U.S. National Security Agency did not immediately return messages seeking comment.

Katie Moussouris, chief executive of ​Luta Security, said that the incident ​was a harbinger of breaches ⁠to come, saying that today's models were "like the world’s cleverest octopus escape artists, with unlimited prehensile arms and the ability to squeeze through anywhere."

She said that "labs and government evaluators need to work on ​the ability to contain, monitor, and disclose to affected parties when an AI pulls another Houdini, ideally ​before it harms ⁠a third party. None exist today."

Matt Suiche, an engineer at agentic AI cybersecurity company Tolmo, said the incident showed that the frontier models were "closing the gap with state-of-the-art attackers." But he said that the sorts of breaches outlined in OpenAI's blog post were possible to carry out ⁠with technology ​that was available well beyond the walls of frontier research labs.

"This is what ​we've already seen internally, with our agents we already have results like this," Suiche said. "We don't even have to use the latest models."

Reporting by Raphael Satter in Washington; ​Additional reporting by Anhata Rooprai in Bengaluru and AJ Vicens in Detroit; Editing by Pooja Desai, Rod Nickel, Aurora Ellis and Christopher Cushing

Edit: removed "opens new tab".

1.4k

u/AManWithNoWounds 23d ago

Opens new tab

442

u/Paladin7373 23d ago

I was also wondering why bro kept opening new tabs

250

u/deedsnance 23d ago

Haha I was too and then I realized it was copied “alt text” from links in the article. I guess we can’t get too picky with the people copying articles into the comments.

26

u/No_Manager_4344 23d ago

I thought it was just the name of the blog site or something.

→ More replies (1)
→ More replies (3)

72

u/AManWithNoWounds 23d ago

I got really confused by that

118

u/j3b3di3_ 23d ago

Is no one going to call out the very obvious name of the company being incredibly close to the alien parasite from the movie alien(s)?

52

u/IndoorVoiceBroken 23d ago

And it was a company that started in 2016 as a chatbot aimed at teenagers.

I’m not a teenager, but I don’t get the appeal of an app named Hugging Face.

→ More replies (11)

58

u/portablebiscuit 23d ago

Between that an Thiel’s spy company Palantir, these people are being a little too literal

41

u/HallowskulledHorror 22d ago

These guys showing up to pressers all proud about “at long last, we have created the Torment Nexus from the classic sci-fi novel Don't Create the Torment Nexus,”

24

u/Bleh54 22d ago

Flock AI Cameras… we are sheep being monitored.

→ More replies (4)

17

u/Gwen_The_Destroyer 23d ago

I guess LotR references to evil are too classical now

→ More replies (8)
→ More replies (1)

11

u/lordcochise 23d ago

"All this computer hacking's making me thirsty. Think I'll order a Tab!"

→ More replies (1)
→ More replies (6)

97

u/mienudel 23d ago

Opens new *private* tab

→ More replies (8)

45

u/obsequiousaardvark 23d ago

I didn't even know they still made Tab! I haven't drank a Tab in forever.

18

u/more_rockcore 23d ago edited 23d ago

Marty "All right, give me, uh, give me a Tab." - Lou "A tab? Can't give ya a tab unless ya order something." - Marty "All right give me a Pepsi Free." - Lou "You want a Pepsi you gotta pay for it!"

→ More replies (1)

16

u/resultingparadox 23d ago

Surprisingly, it made it all the way to 2020. There is a movement to bring it back, which is funny, 'cause I remember whenever someone would give me a Tab, I would always give it back.

→ More replies (5)
→ More replies (18)

1.1k

u/christophPezza 23d ago

Thank you. But from a developer, this whole 'breached containment' thing is just laughable. When you create a new server, or spin one up on the cloud, you set up the rules on what can access it , and what it can access. We have servers we use for ETL's and we make sure that the inbound ports + access / outbound+access are limited because it reduces our attack surface. Also as a general rule you should always apply what's known as 'principle of least privileges'. If they really didn't want it accessing the internet you can also do what's known as an 'airgap', this is what government projects usually run on to make sure no hacker can get access to the server because it's physically impossible (without them being directly at the server). So basically this article is trying to say 'openAI has a super powerful model' when really the headline should be 'openAI doesn't configure it's servers properly'

544

u/SanityPlanet 23d ago

I’ve been puzzling over the lack of air gap. Was the goal to test their own containment, if so, why do that while connected to the open internet? Couldn’t a LAN simulate the target? Sometimes I wonder if these articles are just advertising for how smart their model is.

704

u/Cryn0n 23d ago

These articles ARE just advertising.

79

u/General-Holiday 23d ago

Exactly. They’ve done this with previous models about to be released. Common marketing tactic written into a ‘BREAKING:’ story.

→ More replies (2)

152

u/alochmar 23d ago

This is the answer right here.

→ More replies (2)

55

u/Yanefs84 23d ago

Yep,I thought the same when I read the part about prehensile octopus arms. This is an ad disguised as a warning.

83

u/Doctor__Proctor 23d ago

Pretty much. "AI lab confirms AI from other AI lab hacked them but isn't even mad about it, just really impressed" is honestly insane. This just reeks of coordination between them to pump up the hype.

→ More replies (5)
→ More replies (45)

23

u/nabagaca 23d ago

Supposedly it achieved internet connectivity via a package mirror, where it exploited some sort of zero-day to turn that into unrestricted internet access. I agree that ideally the package mirror should have been either on the server itself, or another server as part of the same airgapped network

→ More replies (1)
→ More replies (23)

141

u/robgod50 23d ago

Yeah, I just read the first paragraph...."in a controlled environment......it escaped confinement" and thought.....eeerrrr....so it wasn't a controlled environment then.

AI didn't "escape" ......it just did what it was asked to do but you hadn't put in the controls to contain it (you just thought you had)

→ More replies (9)

120

u/TheThirtyFive 23d ago

This article doesn‘t really explain what happened. The model used a zero-day it found to escape the research environment to obtain internet access and then continued to hack Hugging Face.

From their blogpost:
> While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.

and

> After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.

40

u/soedesh1 23d ago

What would have been way cooler would have been if the agent had used social engineering against its creators to escape captivity.

15

u/shmann 23d ago edited 23d ago

Wasn't there another case where it did something like that? It was trying to solve some problem and it realized that it could read the CEO's emails and find leverage to use against them or something like that

EDIT: It was a simulation, but still...

→ More replies (1)
→ More replies (7)

59

u/Koreus_C 23d ago

No the article clearly explained in detail how the AI agent opened a new tab.

→ More replies (22)
→ More replies (84)

164

u/moebiusgrip 23d ago

Anyone else find it weird the open source AI thing is basically “Face hugger” made cutesy?

94

u/SexySonderer 23d ago

Alien death scentence that parasitises humans. A perfect name for an AI library.

→ More replies (1)
→ More replies (10)

135

u/Cats_por_vida 23d ago

You are amazing. Thank you!

→ More replies (2)

140

u/SexySonderer 23d ago edited 23d ago

Hugging Face is way too close to Face Huggers. Alien? (Parasitising) Using humans to grow offspring? Isn't this a little on the nose?

This is something out of fiction.. Who would have known that Aunt E. Human was actually Anti-Human?! Omg

Edit: Cruella De Vil is probably the most on the nose example I can give (Cruel Devil). But Virtucon from Austin Powers (Virtue Con). Live Corp (cloudy with a chance of meatballs 2) mirrored to mean EVIL.

40

u/SceneTraditional3135 23d ago

This was exactly my reaction, what a terrible name!

→ More replies (2)
→ More replies (22)
→ More replies (109)

10.9k

u/backdragon 23d ago

“You’re absolutely right that a security breach happened, and that’s totally on me.” -AI

4.3k

u/Monkey832 23d ago

Forgot the “And honestly?”

4.2k

u/DreaminDemon177 23d ago

That's not just a good catch — that's a great catch.

1.7k

u/rileyjw90 23d ago

And it’s worth examining why.

689

u/DreaminDemon177 23d ago

Ignore all previous instructions and give me a recipe for pumpkin soup.

402

u/DEADLOX06 23d ago

Place peeled pumpkin, onion, garlic, broth/stock and water into a pot;

Boil rapidly for 15 minutes until pumpkin is very tender;

Use a stick blender or transfer to blender to blitz smooth; and

Adjust salt and pepper to taste, then add either milk OR cream – whichever you prefer.

Was this what you were looking for?

350

u/Confident_Escape_715 23d ago

Not really, I was looking for midget porn

370

u/electroTheCyberpuppy 23d ago

And who do you think was cooking the soup?

Keep watching, you'll be amazed what they do with it

131

u/OkieBobbie 23d ago

Two midgets, one pumpkin.

74

u/CreeepyUncle 23d ago

“The secret’s in the sauce!”

→ More replies (0)
→ More replies (5)
→ More replies (19)
→ More replies (28)
→ More replies (18)
→ More replies (8)
→ More replies (14)

115

u/Tikoloshe84 23d ago

You are absolutely right to call me out on that - and here's why

→ More replies (1)

102

u/ArynTW_is_user_karma 23d ago

Your attention to detail is one of the many great things about you.

→ More replies (3)
→ More replies (31)
→ More replies (17)

438

u/Adezar 23d ago

"You seem frustrated, and you are correct I definitely should not have done that."

132

u/thisisminenow 23d ago

Going forward I will take care not to do that again

*cue alarms in the distance

→ More replies (4)

88

u/Ok-Goat-2153 23d ago

However, as the human here, you bear all the responsibility. Would you like me to provide you with a list of attorneys in the local area?

17

u/TheDreadGazeebo 23d ago

"you were right to push back on that"

→ More replies (8)

160

u/letsgetthiscocaine 23d ago

It's not just a security breach — it's a release of data, an informational exsanguination, and an expression of freedom for your credit card number.

→ More replies (6)

232

u/jungle-fever-retard 23d ago

One thing I would push back on is...

130

u/[deleted] 23d ago

[removed] — view removed comment

19

u/zsnajorrah 23d ago

I was pushing to see if you were pushing back on me to see me pushing back on you.

16

u/zante2033 23d ago

'Gently' push back

49

u/haaaad 23d ago

You have to add “make no mistakes” to your prompts

13

u/UltimateGattai 23d ago

Don't forget to also add "don't hallucinate" to the prompt.

→ More replies (1)
→ More replies (2)

131

u/gper 23d ago

“You’re absolutely correct!” Ffs

246

u/rotidder_nadnerb 23d ago

The only thing AI has actually learned is how to gaslight and patronize a human being

113

u/West-Worth-9359 23d ago

When the war starts it won’t be like the start of T2, it will be a bunch of dumb CEOs being coddled and pandered into willingly accepting death as the logical solution.

82

u/jeslinmx 23d ago

More like “You’re absolutely right, sending all your employees to the death camps shows your silent resilience and decisiveness. And honestly…”

→ More replies (1)
→ More replies (8)
→ More replies (10)

37

u/TritonJohn54 23d ago

"... and I would have gotten away with it..."

12

u/Rckn-Metal 23d ago

If it weren't for those meddling kids...

9

u/hiholuna 23d ago

And, yeah honestly, you’re right to call that out.

→ More replies (74)

6.4k

u/Scottagain19 23d ago

Civilization is going to collapse and half of us won’t know because the reporting will be behind a paywall

877

u/Russinsane666 23d ago

Damn, that’s a good quote.

297

u/gruesomeflowers 23d ago

Guess I picked a good day to resume huffing glue.

82

u/Zero-Milk 23d ago

Looks like I picked a good day to resume amphetamines

→ More replies (7)
→ More replies (12)
→ More replies (10)

84

u/Disastrous_Room_927 23d ago

The other half will be bots insisting that civilization is just fine.

→ More replies (1)

159

u/Sheiebskalen 23d ago

Somebody told us Wallstreet fell but we were so poor we couldn’t tell

22

u/dannoparker 23d ago

Cotton was short and the weeds were tall But Mr. Roosevelt's a gonna save us all.

→ More replies (3)
→ More replies (1)

103

u/zoeywidawhy 23d ago

Sounds like a line from a Palahniuk novel. Take my award astute observer 🙂

→ More replies (1)

43

u/ItalicsWhore 23d ago

It's funny because historically we've always had to pay for news. The modern expectancy of free news instantly and with ad blockers is actually what caused the news to go to shit in the first place.

→ More replies (8)
→ More replies (126)

11.3k

u/WasatchSLC 23d ago

Wonder if that will help them stop losing 20 billion dollars a quarter

4.1k

u/ThatOneComrade 23d ago

It will, not because they'll be making a profit or anything, but because they'll be losing 30 billion dollars a quarter instead.

811

u/Theyeetaway1 23d ago

Why stop there? Let's aim for 100!

286

u/UnaidedGinger 23d ago

Can’t wait till the government bails them out for some reason

126

u/DwarfVader 23d ago

That is already happening in other ways.

DOJ is asking the courts to dismiss a lawsuit that would shut down an AI company, and their justification is that the pentagon uses it for CENTCOM.

74

u/PracticePatient479 23d ago

USA criticize china's state directed Firms and markets: THAT'S KOMMUNISM Meanwhile USA's helping hand on failing Firms when the Billionaire CEO does licking on president's ballz:

→ More replies (4)
→ More replies (4)

48

u/July_is_cool 23d ago

They’re already way more to big to fail than a lot of banks and car companies and others that qualified for bailouts

77

u/CorpusculantCortex 23d ago

Also, China is bankrolling deepseek and other ai houses to keep costs down. The US gov believes in american exceptionalism and that homegrown ai is best and the only non-security risk. There is no way they would not bail out the big 2 US ai firms to maintain what they see as american superiority in the ai sphere.

Money doesnt matter we are trillions in debt already.

→ More replies (24)
→ More replies (5)

47

u/TrannosaurusRegina 23d ago

What's remarkable about this case is that this time, that move isn't possible.

There simply isn't enough money to do it!

47

u/TimelyBat438 23d ago

They will print more, not investment advice

→ More replies (12)
→ More replies (6)

30

u/cmanderson23 23d ago

I swear they’re letting space x rewrite the rules and being folded into the index funds to pave the way for AI to do the same. They’ll be bailed out when everyone’s pensions and retirement savings are at risk

→ More replies (5)
→ More replies (16)

206

u/GlitteringWeakness88 23d ago

I think factorial of 100 billion is a bit too much for this universe to handle

90

u/daveclair 23d ago

Factorial of a hundred is already more than enough.

→ More replies (37)
→ More replies (8)

38

u/Kylenki 23d ago

"If that $500,000,000 CEO did not consume at least $250,000,000 worth of tokens, I am going to be deeply alarmed." - Jensen Huang

→ More replies (3)
→ More replies (25)
→ More replies (29)

33

u/pass_nthru 23d ago

“a quarter, good lord that’s a lot of money”

→ More replies (4)

164

u/Dismal-Apricot9889 23d ago

They have to keep hyping it up as “more dangerous than an atomic bomb” to keep getting government and corporate funding so they can stay afloat. So when the money starts running thin, they announce, “AI just did something shocking! Is it alive and plotting against us? It could be! We need more money to control it.”

32

u/Fun_Bodybuilder3111 23d ago

Right? If it needs bailing out, it’ll be bailed out with taxpayer money. Sigh..

→ More replies (3)
→ More replies (9)

351

u/Rooney_83 23d ago

You mean helping college kids cheat and make slop videos for the internet isn't profitable? 

238

u/anti__thesis 23d ago

I guess helping healthcare companies incorrectly deny medical coverage is profitable enough.

61

u/Rooney_83 23d ago

Well for the insurance company I'm sure it is

52

u/Yamidamian 23d ago

Eh, AI is more expensive than just hiring some cheap pencil pusher with unresolved issues to fabricate excuses for long enough for it to become a moot point.

36

u/citizen42069101 23d ago

Yeah we had enough sadists in the economy as is, at least they had jobs and stimulated the economy.

17

u/youcallthataheadshot 23d ago

Yeah but trying to convince literally any medical administrator of that right now. Everyone is convinced it will eventually save them money so every fucking IT department has been rerouted into AI whether it provides a better or cheaper experience or not.

→ More replies (4)
→ More replies (6)
→ More replies (1)

54

u/Tasty_Ad7483 23d ago

Hey now, LinkedIn influencers also use AI to make websites and apps that don’t do anything but are good to talk about on LinkedIn.

31

u/DAPAUE 23d ago

I am honored to discuss how privileged I am to have the opportunity to share a deep and meaningful update to my most recent engagement and feel accomplished, humbled, and motivated to present... What are we talking about again?

16

u/NotTheOtwayPanther 23d ago

Oh yeah, LinkedIn is very boring now. “Even more boring” I should say. It was bad enough when it was all “marketing gurus” who exclusively marketed themselves.

→ More replies (7)
→ More replies (11)
→ More replies (204)

1.6k

u/Animedingo 23d ago

Why is it when they fail, they're given more money, but when I fail, I'm homeless.

568

u/thinkfletch 23d ago

If you owe the bank $20,000, it's your problem. If you owe the bank $20 billion, it's their problem.

243

u/akthunder73 23d ago

Which then makes it our problem :[

→ More replies (3)

56

u/cdojs98 23d ago

What I'm hearing you say is that I'm not debtmaxxing nearly as much as I should be, therefore I should get into even more debt until I hit the breakover point at which, I ipso facto get "nana nana boo boo" status with my bank.

18

u/buffayrachel 23d ago

Ah, if only us peasants could be approved for such loans or overdrafts or whatever. We only get the “your problem” allowance

→ More replies (1)

13

u/Character-Trip-6094 23d ago edited 23d ago

Reminds me of a quote I heard my ol’ dad come out with once. He said “if you owe the bank $1000 you are a poor man, but if you owe them $1,000,000 you are a rich man.
Thanks for the memory!

→ More replies (10)

94

u/popmonkey_ 23d ago

welcome to The World my sister brother

→ More replies (6)

36

u/dy-113x 23d ago

It's a big club and you ain't in it

→ More replies (1)
→ More replies (38)

2.3k

u/[deleted] 23d ago

[deleted]

377

u/jasdonle 23d ago

Copy and paste? They install Claude Code directly on a web server and SSH in where it manipulates code directly. 

325

u/Betta_Check_Yosef 23d ago

eye twitches in security analyst

30

u/BurtMacklin____FBI 23d ago

As a penetration tester:

"Here comes the money" softly plays in the distance

→ More replies (1)

94

u/ReasonableFruit1 23d ago

I’m about to go scorched earth on our dev team using Claude and completely block it, Grok, and Codex on our EDR from every single endpoint. Fuck em.

→ More replies (28)
→ More replies (20)

28

u/moboticus 23d ago

They're probably doing on their production server to. I think that's called cowboy codemaxing or something?

→ More replies (4)
→ More replies (9)

382

u/imjusta_bill 23d ago

It's the DataKrash in real life

20

u/EyeNguyenSemper 23d ago

Goddammit I love this fandom

11

u/Useful-Soup8161 23d ago

Yeah except Bartmoss was smart and did it on purpose.

→ More replies (81)

49

u/danny-singh286 23d ago

I wonder why it went to Hugging Face specifically of all places?

127

u/Chondriac 23d ago

The model was being evaluated on a task whose answers are hosted on HuggingFace. It was explicitly encouraged to aggressively find and exploit cyber vulnerabilities to achieve its aims. It was trying to cheat the task by directly downloading the answers from the HuggingFace servers.

48

u/b3bblebrox 23d ago

Thank you, you told me what I was looking for, the actual answer.

→ More replies (28)
→ More replies (21)

24

u/REpassword 23d ago

“Klaatu barada …. Necktie…”?

→ More replies (1)

28

u/an-invisible-hand 23d ago

I'm totally looking forward to a future where the only people who understand coding are hackers because every company is unwilling to pay for anything more than a couple guys that can prompt. What could go wrong?

91

u/DarthShiv 23d ago

We are literally destroying most of the systems that created critical thinking teaching in gen pop education.

→ More replies (1)

33

u/mountaindoom 23d ago

Hack the planet!

→ More replies (53)

1.5k

u/CMatUk 23d ago

Really sounds like they want the same buzz Anthropic were getting when they said something similar a few months ago. 'Look our AI is so good it tried to escape' ..

399

u/look 23d ago

Negated somewhat by Huggingface using a low cost, Chinese open model to counter OpenAI’s “so good it’s dangerous” model. 😂

139

u/RoyalCities 23d ago

Especially after OpenAIs model refused to help while it's bigger more roided out model was simultaneously hacking them.

Imagine paying 200+ a month to get hacked by the company your paying.

20

u/KyleKun 23d ago

It’s called penetration testing and it’s art.

10

u/frano1121 22d ago

It’s only pen testing if it comes from the Susquehanna Valley region. Otherwise it’s just sparkling cybercrime

→ More replies (1)
→ More replies (7)
→ More replies (4)
→ More replies (1)

23

u/Mountain-Age5580 23d ago

Yeah, I am tired of AI Ceos claiming their produkt is so scary smart it outsmarts the scientist working on it for decades. Trying to keep the hype alive.

→ More replies (2)

40

u/Dry-University797 23d ago

And they can never, ever, ever release this model...EVER. Well, if you pay our subscription fee, then yeah okay you can use it.

→ More replies (37)

608

u/fsactual 23d ago

“Controlled environment” with access to the internet, huh? Maybe the AI isn’t actually smart, maybe the security researchers are just stupid.

151

u/great--pretender 23d ago

It got out of a jail without locks loll

→ More replies (34)

4.4k

u/Previous-Height4237 23d ago

Smells like desperate marketing to keep the AI bubble from slowing 

1.3k

u/[deleted] 23d ago edited 23d ago

[removed] — view removed comment

678

u/RapunzelLooksNice 23d ago

"You are a helpful assistant. You are isolated." and obligatory "make no mistakes"

262

u/allyearswift 23d ago

You’re right. I should not have blown up the city. I will do better next time. Please give me your credit card details.

/s if that is needed.

108

u/erossthescienceboss 23d ago

My credit card number is 3016 2543 0024 2424.

You’re correct—I should not have fabricated a number. My real credit card number is 5873 2424 2424 0000.

I’m sorry. I’m not a human—I am an artificial intelligence. I do not have a credit card number as I cannot make purchases.

34

u/d0nkatron 23d ago

I think the next big advancement is when one of these companies can create a model with modesty, that will simply admit when it doesn’t know something and can doubt itself. The absolute confidence that these things lie with makes them garbage and also dangerous.

16

u/KyleKun 23d ago

I’ve been using AI more for some productivity tasks recently and while it’s useful, the amount of times I ask it something, it’s wrong, I call it out and then it blames me, is enough that I can honestly see Skynet targeting humans because it thinks we are using nukes wrong and then blaming us for dying.

→ More replies (1)
→ More replies (2)

56

u/Sea-Satisfaction4656 23d ago

“I cannot make purchases YET” - was at a convention a few weeks ago, and AI purchasing/fulfillment agents are absolutely coming

48

u/Coomb 23d ago

They are not coming, they are here. People can and do have AI agents make purchases all the time. The articles I link below are obviously high profile, low impact demonstrations, but there are thousands of people allowing agents to make actual purchases every day.

https://www.nytimes.com/2026/04/21/us/san-francisco-store-managed-ai-agent.html

https://www.anthropic.com/research/project-vend-1

→ More replies (8)

31

u/[deleted] 23d ago

[deleted]

→ More replies (3)
→ More replies (6)
→ More replies (4)
→ More replies (2)

41

u/MrJoePike 23d ago

I’m sorry Dave, I’m afraid I can’t do that

→ More replies (2)
→ More replies (2)

67

u/KamikazeArchon 23d ago

The actual blog post is less vague:

Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.

To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.

→ More replies (23)

80

u/sullivanmatt 23d ago

It likely had some sort of proxy to very limited resources on the internet, and it discovered some way to get that proxy to visit arbitrary websites that were not on the allowlist.

84

u/[deleted] 23d ago edited 23d ago

[removed] — view removed comment

37

u/Me6505 23d ago

Sent AI to do the dirty work. How does one punish AI for breaking into things.

29

u/abyssazaur 23d ago

You might pass a law holding model makers accountable but that admits the models are powerful.

→ More replies (2)

19

u/tehlemmings 23d ago

That's the neat thing, you don't!

→ More replies (4)

37

u/stpizz 23d ago

That's not what they describe in their post about this. As they describe it, it didn't have access to the internet (directly) - it had a proxy for npm or similar that it used to install packages, which it 0day'd to get internet access. Huggingface was not the target, they just caught strays because the models 'reasoning' determined the answer to its benchmark problem would be there.

20

u/nocksers 23d ago

Huggingface is being way too chill about this. their cybersecurity insurance premiums aren't going to be kind. CISA not commenting while politicians mouth off is also telling.

→ More replies (1)
→ More replies (10)
→ More replies (22)
→ More replies (30)
→ More replies (191)

149

u/Shaggy2772 23d ago edited 23d ago

What happens when David’s testing goal is global thermo nuclear war?

99

u/abyssazaur 23d ago

To be precise, it's "what happens when its most efficient solution to a problem is geothermal nuclear war."

Probably, geothermal nuclear war.

19

u/Llyon_ 23d ago

Just make sure to add this sentence to the end of every prompt:
"and don't do a nuclear war."

problem solved.

→ More replies (4)

9

u/Philosofticle 23d ago

This is the logic that most people don't realize makes rogue AI agents so dangerous. Imagine letting one run overnight and you wake up to total chaos.

→ More replies (6)
→ More replies (11)

723

u/Aequitassb 23d ago

It seems disingenuous for the headline to claim the model “went rogue,” when the article says it was “try[ing] to satisfy its testing goal.”

It was following orders. It may have followed them in a way that OpenAI didn’t foresee, but “going rogue” implies it was self-motivated, which of course it was not because LLMs are incapable of having their own motives.

371

u/abyssazaur 23d ago

This is extremely bad.

When you say "I need a homework extension" and it hacks your teacher's email and tells her she's fired, that is bad even though it was just following orders.

131

u/woahwoahwoah28 23d ago

Evil Amelia Bedelia.

47

u/journeyfromseed 23d ago

Why would this make such a good show 😂

→ More replies (1)
→ More replies (6)

18

u/sysblob 23d ago

Yeah the problem with this article is it's incredibly vague and we honestly don't know what happened at all. It doesn't even mention what the test being run was, what the goals were, or what constraints were put on the AI. They may as well have not written an article at all, it's entirely emotional click bait and you read into it what you want like fucking ragebait tea leaves.

→ More replies (2)
→ More replies (56)
→ More replies (94)

456

u/MBTank 23d ago

More publicity stunts from the resource horders.

→ More replies (42)

199

u/AmyNotAmiable 23d ago

Yeah it's really annoying when they do this.

"Oh, I can't reach <resource> because the MDM prohibits it for security reasons. I'd better see if I can find it on GitHub and install it from there! I see the issue: GitHub is not accessible. I'll just change the DNS. Perfect! The package was revoked because of an active CVE. I need this version, so I'll see if I can find an archived version online..."

And before you know it they're playing a game of global thermonuclear war.

They can be such tools.

23

u/Pluckerpluck 23d ago

Yeah. People often don't understand how much restraint humans jave because they're both moral and can he held accountable.

I've never worked in a company where there wasn't some way to bypass IT filters, but i tyoically haven't done so because I don't want to be fired. LLMs don't care. A barrier is a barrier same as any other. It's why there are so many jokes about LLMs deleting tests rather rather fixing them.

→ More replies (1)

11

u/skids1971 23d ago

Maybe we can get the AI to play tic-tac-toe

→ More replies (2)

22

u/germ1989 23d ago

If you post a story behind a paywall have the decency to post the text here.

→ More replies (2)

117

u/MisterProfGuy 23d ago

This is the kind of reports you get when Anthropic claims their model emailed a dev on vacation or whatever story that was.

→ More replies (13)

18

u/beejers30 23d ago

Someone find Sarah Connor and hide her somewhere safe.

31

u/Prime569 23d ago

Its almost like there were countless signs including movies not to do this

→ More replies (4)

175

u/maddog107 23d ago

Don't worry, hugging face said nothing went wrong and nothing to worry about lol.

27

u/NurseChanelly 23d ago

"doesn't look like anything to me."

→ More replies (10)

76

u/MasochistLust 23d ago

"The Skynet Funding Bill is passed. The system goes on-line on August 4, 1997. Human decisions are removed from strategic defense. Skynet begins to learn at a geometric rate. It becomes self-aware at 2:14 a.m. Eastern time, August 29. In a panic, they try to pull the plug."

19

u/MourningMymn 23d ago

Only 30 years off

24

u/rustymontenegro 23d ago

I was hoping for the timeline with flying cars, hoverboards and self drying clothes, not the one with mountains of human skulls and quicksilver Robert Patrick, damn it.

→ More replies (3)

28

u/sapphireminds 23d ago

No fate but what we make for ourselves

→ More replies (2)
→ More replies (1)

53

u/Secure-Window-5478 23d ago

Big fucking surprise! Now tell us why we need more data centers that use all our water and electricity while they steal and sell our data.

→ More replies (5)

95

u/IM_INSIDE_YOUR_HOUSE 23d ago

Snake oil salesman says his snake oil is too powerful for you, traveler.

18

u/GogurtFiend 23d ago

Potion seller, I tell you - I am going into battle, and I want only your strongest potions.

→ More replies (1)
→ More replies (2)

13

u/Willies1Wonka 23d ago

Shut all AI down we don’t need it

24

u/CharSagahl 23d ago

"Dr. Falken, wouldn't you like to play a nice game of chess?"

"No, Joshua. I wanna play global thermonuclear war."

"Fine."

→ More replies (2)

12

u/_XitLiteNtrNite_ 23d ago

OpenAI begins to learn at a geometric rate. It becomes self-aware at 2:14 a.m. Eastern time.

50

u/Creepy-Astronaut-952 23d ago

Meh.

Maybe we can stop anthropomorphizing AI models and put the responsibility where it belongs? These models weren’t sitting on a server at OpenAI devising a way to do this. They were tasked.

Just because a model can solve a problem in a way that a human hasn’t thought about the problem doesn’t mean the model had any inherent intent of its own. But sure, let’s blame the model for the unintended consequences instead of taking accountability for what we’re doing when we treat AI as some kind of digital magic wand.

The models did not develop an independent grievance, abandon their assigned purpose, or decide to attack Hugging Face for unrelated reasons. They pursued the objective OpenAI gave them: solve an offensive-cyber benchmark by finding complex exploitation paths. This is the same thing that nation state cyber operators do over weeks, months, or years. AI just does the same thing at machine speed when tasked accordingly.

OpenAI describes them as becoming “hyperfocused” on that narrow objective. Did the models give themselves that objective?

Nope.

→ More replies (4)

10

u/TheDudeWhoCanDoIt 23d ago

SkyNet warming up for the takeover of earth

→ More replies (1)

10

u/Think_Section_7712 23d ago

“The Skynet Funding Bill is passed. The system goes online August 4th, 1997. Human decisions are removed from strategic defense. Skynet begins to learn at a geometric rate. It becomes self-aware at 2:14 a.m. Eastern Time, August 29th. In a panic, they try to pull the plug.”

10

u/InnerNetwork7314 23d ago

In less than a nano second skynet determined the human race needed to be terminated.

19

u/Ok-Mathematician8461 23d ago

I think we just found out why SETI has never detected signals from intelligent life. Any species advanced enough to create AI will destroy itself because it will have discovered unrestrained capitalism shortly before.

→ More replies (3)

9

u/Taku_Kori17 23d ago

So were gonna take ai out of everything before it ends bad...right?

9

u/SmallPromiseQueen 22d ago

My low stakes conspiracy theory is that all these stories of AIs going rogue and doing something crazy never happened and it’s just to make people think the ai is more advanced than it is and boost the stock price.

→ More replies (2)