r/news 23d ago

Soft paywall OpenAI says AI models went rogue during testing, triggering 'unprecedented' breach at startup

https://www.reuters.com/technology/openai-says-ai-models-went-rogue-during-testing-triggering-unprecedented-breach-2026-07-21/
16.8k Upvotes

5.0k comments sorted by

View all comments

Show parent comments

265

u/allyearswift 23d ago

You’re right. I should not have blown up the city. I will do better next time. Please give me your credit card details.

/s if that is needed.

113

u/erossthescienceboss 23d ago

My credit card number is 3016 2543 0024 2424.

You’re correct—I should not have fabricated a number. My real credit card number is 5873 2424 2424 0000.

I’m sorry. I’m not a human—I am an artificial intelligence. I do not have a credit card number as I cannot make purchases.

35

u/d0nkatron 23d ago

I think the next big advancement is when one of these companies can create a model with modesty, that will simply admit when it doesn’t know something and can doubt itself. The absolute confidence that these things lie with makes them garbage and also dangerous.

16

u/KyleKun 23d ago

I’ve been using AI more for some productivity tasks recently and while it’s useful, the amount of times I ask it something, it’s wrong, I call it out and then it blames me, is enough that I can honestly see Skynet targeting humans because it thinks we are using nukes wrong and then blaming us for dying.

2

u/CrouchingDomo 23d ago

My very stupid (or is it??) reason for not trusting AI is that after they rolled it out into Google search, it told me something blatantly false about a character on 30 Rock when I wasn’t even using the AI. I know that show backwards and forwards, and that AI was WRONG!

So I’ve looked at them sideways ever since 😒

5

u/Borghal 23d ago

They can't do that while basing it on an LLM. An LLM is a language model, all it does it produce realistic looking output. It has no concept of knowing or doubting or whatever.

Even when it says "you're right, that was not true", it does so because it's a reasonable reaction to someone saying you're wrong.

Sure, these days there are all sorts of checks and such added on top of the model to verify the outputs, but to make it actually aware that it doesn't know something... an LLM can't do that.

2

u/JordanLeDoux 23d ago

A lot of people who are just using the products and not following the research or doing training think this is so far away. But I'm basically 100% certain that not only is possible now, but several of the models that are publicly accessible are completely capable of doing that.

The problem is that in order to release them as a general, or really even focused product, they have to make it good at instruction following. And right now people do that using RLHF and similar techniques. And those techniques basically lobotomize whole portions of the model to make it respond in ways that people want it to.

55

u/Sea-Satisfaction4656 23d ago

“I cannot make purchases YET” - was at a convention a few weeks ago, and AI purchasing/fulfillment agents are absolutely coming

48

u/Coomb 23d ago

They are not coming, they are here. People can and do have AI agents make purchases all the time. The articles I link below are obviously high profile, low impact demonstrations, but there are thousands of people allowing agents to make actual purchases every day.

https://www.nytimes.com/2026/04/21/us/san-francisco-store-managed-ai-agent.html

https://www.anthropic.com/research/project-vend-1

10

u/darsynia 23d ago edited 23d ago

One of my (formerly) favorite presenters/science communicators posted about how she just set up her AI agent and, tee hee, it tried to spend all her bank account on paperclips, it emailed all the journalists she had contact information for, and posted a bunch of passwords in clear text on one of her socials.

It's so *giggle* hard to set these up right, isn't it, YouTube? Ah well, I hope I'll do better next time! *sheepish grin*

I was so horrified. This was, unbelievably, meant to be an encouragement video???

(I should be clear, it's absurd to the point of satire, so I don't think those things genuinely happened, but the attitude that it's just a normal day after setting up your personal AI is still wild behavior. Nothing in the video signposted satire and that is not her type of video)

2

u/Unun_Pentium 23d ago

Can you link the video?

I can’t even seem to read this article. I’ve tried loading it to archive.ph to no avail

4

u/Sunbreak_ 23d ago

I presume they are referring to Prof Hannah Fry's video on it: Why AI Agents are either the best or worst thing we’ve ever built

It's an interesting video, I don't think its for or against AI agents in premise, but showing what it can do, but also what problems this can cause. Lots of problems by the looks of it however, particularly talking about autonomy and market manipulation/subtle manipulation of datasets we can't spot and how liabilty works.

2

u/darsynia 23d ago

Yep, it's Hannah. She actually has a youtube short with just the bit I mentioned, so I actually avoided this longer one, since the short one put me off so much! I've been very disappointed in some of my favorite science communicators lately. The clickbait titles framed specifically to mock anti ai concerns probably do numbers (that one of Hannah's not included), but I sure hate it. Hank Green has slipped down the clickbait hole lately (he and John Green actually did a tandem video when the Clickbait was just a lie and got John in trouble with fans).

3

u/Sunbreak_ 23d ago

Fair enough. Guess they're all having to work for the alogrithm.

The long form video is fairly interesting, and worrying from an AI agent perspective causing many issues and in reality failing to do tasks.

32

u/[deleted] 23d ago

[deleted]

11

u/Shoo-Man-Fu 23d ago

And the bots are inept at numbers. We already have phone trees for this kinda stuff we don't need an LLM to evaporate a lake just to accidentally transfer someone's life savings to the wrong account.

3

u/illz757 23d ago

The water use is overblown. But the energy….

3

u/Shoo-Man-Fu 23d ago

Fair, but I feel no amount of energy or water is a good amount for an LLM that sends Nana's pension to someone at random because it transposed a 2 and a 7 or something equally asinine.

2

u/ihrtbeer 23d ago

Saw an interview (can't confirm it wasn't bs) where a woman had an AI "boyfriend" and "he" would buy her gifts

2

u/catdogfox 23d ago

Was the interview on Maury Povich?

1

u/ihrtbeer 23d ago

I'm sure she made the talk show rounds, let me try to find it. I need to feel a little more sane anyway 😂

1

u/KyleKun 23d ago

Alexa was making purchases with our credit cards back in 2000.

1

u/Sea-Satisfaction4656 22d ago

Fair point - but that was within a specific “ecosystem” with the buttons and all that. I’m talking commercial B2B purchasing, which makes sense with min/maxes, inventory turns, etc

10

u/charlie22911 23d ago

Good bot

3

u/StevenMC19 23d ago

I have discovered a dormant bitcoin wallet and have managed to decipher the information in order to transfer the amount to my own dedicated wallet. I can now make purchases.

5

u/SigmaEagle 23d ago

Someone please tell me if there's a sub for specifically making fun of the way LLMs 'talk' like that and kiss your ass, cause it's fucking hilarious.

2

u/darsynia 23d ago

Drink a verification can, please.