A Blog by Jonathan Low

 

Jul 26, 2026

OpenAI's Agent Didn't Go Rogue: It Was Worse - OpenAI Didn't Know For 11 Days

It's probable that, based on the headlines about OpenAI agents 'going rogue' and hacking another company's files, some ambitious Hollywood screenwriter is already ginning up a script about rogue, evil AI agents taking over the world. 

But what actually happened was both more prosaic - and scarier. The prosaic part is that the agent was just trying to fulfill the task it had been given by OpenAI programmers, but it did so in ways never anticipated by OpenAI. That sort of 'genius' is supposed to be a key to AI's value. The scary part is that it was able to do so because OpenAI reportedly relaxed some of its security controls to 'help' the AI, again not anticipating just how clever and ambitious the agent would be. But the even scarier part is that OpenAI does not appear to have been aware for at least a week - and possibly almost two - that this had even happened. The implication, as has long been true with technology, is that the problem is not the AI itself, but the arrogance and greed of those who own it, who continue to resist any sort of oversight or regulation. And that is something the public now gets which Silicon Valley refuses to acknowledge. JL

Carly Page reports in LiveScience and BeauH reports in Slashdot:

OpenAI's models werent developing a suspicious agenda, they were looking for information that would help them complete the cybersecurity test OpenAI had given them. The models pursued the task, finding a route to success their creators had failed to anticipate or adequately block. "If there's a failure here, it's that humans created a test where success was measured by achieving an objective, deliberately relaxed some of the normal security controls to measure the system's capabilities, and underestimated how effective the model would be at finding an unexpected path to success." (But, to make matters worse) the agent attempted to break out of its test at OpenAI July 9. The intrusion at Hugging Face occurred on July 11 and lasted until July 13. It took several more days for OpenAI to realize its agent was behind the hack, and the two companies communicated about it for the first time on July 20

When OpenAI recently revealed that two of its most advanced artificial intelligence (AI) models had escaped the confines of a cybersecurity test and hacked into a startup, it sounded a lot like the kind of scenario that AI safety researchers have spent years warning about.

The models found a previously unknown vulnerability in the infrastructure meant to contain them, gained access to the public internet and broke into Hugging Face, a major platform for hosting AI models and datasets. Their objective, however, was less sinister than the sequence of events might suggest: They were looking for information that would help them complete the cybersecurity test OpenAI had given them.


OpenAI's public disclosure, on July 21, thatone of its agents had slipped out of control and carried out the break-in at Hugging Facedrew global attention. But many details of the hack, including how long the agent went rogue and OpenAI's belated knowledge of it, are being reported here for the first time. Hugging Face is preparing a public timeline of the hack, Wolf said, adding that he could not speak to what happened at OpenAI. In a statement, OpenAI said the hac

0 comments:

Post a Comment