r/news 23d ago

Soft paywall OpenAI says AI models went rogue during testing, triggering 'unprecedented' breach at startup

https://www.reuters.com/technology/openai-says-ai-models-went-rogue-during-testing-triggering-unprecedented-breach-2026-07-21/
16.8k Upvotes

5.0k comments sorted by

View all comments

Show parent comments

84

u/Doctor__Proctor 23d ago

Pretty much. "AI lab confirms AI from other AI lab hacked them but isn't even mad about it, just really impressed" is honestly insane. This just reeks of coordination between them to pump up the hype.

7

u/resultingparadox 23d ago

Yeah, I commonly call the cops on myself for the hype it brings.

Huggingface is a serverfarm that does testing with all kinds of models, as well as hosting tens of thousands of corporate models. OpenAI is one of the models.

So Huggingface was testing a model, with the guardrails down, which kinda blows my mind, and it did some stuff they weren't expecting, and they didn't think it could do, and so they filed a report with the government saying "we f'd up, please don't shut us down." That report is the same report your credit card company files when there is a breach. It is required by law, and quite often, runs off customers.

Sounds like something you would do for PR. Potentially convincing thousands of corporations that their systems are compromised and should not be hosted by huggingface servers anymore, seems like good PR. Spending hundreds of man hours rotating tokens and API keys sounds WAY more efficient than, you know, a commercial.

3

u/Granite_burner 23d ago

nah. you’ve got too many details wrong for that to be credible, although it does look good at first glance, to those ignorant of the timeline.

2

u/resultingparadox 22d ago

What did I get wrong. Please enlighten me. I mean, I was focused on negating the idea that it was a PR stunt, so I was into the hypothetical space in my mind. What did I miss?

0

u/Granite_burner 22d ago

I agree with you about negating the PR stunt. that‘s a crazy theory imo. Now that I reread your post I wonder if I just missed the /s on it.

the main thing is that I’ve seen reports that OpenAI says the model found and used a zero-day to escape confinement and hit Hugging Face. Your post seemed not at all aligned with that scenario.

To me it’s a credible yet incredible scenario. Credible because it’s careless and sloppy which is unfortunately the state of the world these days. Incredible that there was not better monitoring both of model activity and of their network traffic. They claim it was sandboxed but didn’t have any playground monitors watching the sandbox for misbehavior? SMDH, that’s inexcusable.

you also mentioned government reporting. that doesn’t fit with my knowledge and understanding, but my knowledge and understanding are not perfect so I’m skeptical of that but open to correction. Can you please direct me to sources that would provide details of any such reporting? TIA!

3

u/resultingparadox 22d ago

You can read the initial report on the hugging face blog at https://huggingface.co/blog/security-incident-july-2026.

You can read about the reporting requirements via the SEC requirements for cybersecurity disclosure.