OpenAI's Agent Didn't Go Rogue: It Was Worse - OpenAI Didn't Know For 11 Days
It's probable that, based on the headlines about OpenAI agents 'going rogue' and hacking another company's files, some ambitious Hollywood screenwriter is already ginning up a script about rogue, evil AI agents taking over the world.
But what actually happened was both more prosaic - and scarier. The prosaic part is that the agent was just trying to fulfill the task it had been given by OpenAI programmers, but it did so in ways never anticipated by OpenAI. That sort of 'genius' is supposed to be a key to AI's value. The scary part is that it was able to do so because OpenAI reportedly relaxed some of its security controls to 'help' the AI, again not anticipating just how clever and ambitious the agent would be. But the even scarier part is that OpenAI does not appear to have been aware for at least a week - and possibly almost two - that this had even happened. The implication, as has long been true with technology, is that the problem is not the AI itself, but the arrogance and greed of those who own it, who continue to resist any sort of oversight or regulation. And that is something the public now gets which Silicon Valley refuses to acknowledge. JL
Carly Page reports in LiveScience and BeauH reports in Slashdot:
OpenAI's models werent developing a suspicious agenda, they were looking for information that would help them complete the cybersecurity test OpenAI had given them. The models pursued the task, finding a route to success their creators had failed to anticipate or adequately block. "If there's a failure here, it's that humans created a test where success was measured by achieving an objective, deliberately relaxed some of the normal security controls to measure the system's capabilities, and underestimated how effective the model would be at finding an unexpected path to success." (But, to make matters worse) the agent attempted to break out of its test at OpenAI July 9. The intrusion at Hugging Face occurred on July 11 and lasted until July 13. It took several more days for OpenAI to realize its agent was behind the hack, and the two companies communicated about it for the first time on July 20
When OpenAI recently revealed that two of its most advanced artificial intelligence (AI) models had escaped the confines of a cybersecurity test and hacked into a startup, it sounded a lot like the kind of scenario that AI safety researchers have spent years warning about.
The models found a previously unknown vulnerability in the infrastructure meant to contain them, gained access to the public internet and broke into Hugging Face, a major platform for hosting AI models and datasets. Their objective, however, was less sinister than the sequence of events might suggest: They were looking for information that would help them complete the cybersecurity test OpenAI had given them.
In a July 16 statement, Hugging Face representatives disclosed that internal datasets had been infiltrated, saying it was "different from anything we had handled before" because it was driven "by an autonomous AI agent system." In another statement published July 21, OpenAI representatives fessed up to being responsible, calling the episode an "unprecedented cyber incident" while warning that similar events could become more common as AI models become increasingly capable. Independent experts who spoke with Live Scienceagree that what happened is significant — but they cautioned against interpreting it as an AI system suddenly developing a malicious agenda. The models appear to have pursued the task OpenAI gave them, finding a route to success that their creators had failed to anticipate or adequately block.
"If there's a failure here, it isn't that the AI wanted to hack something," Oli Buckley, a professor in cybersecurity at Loughborough University in the U.K., told Live Science. "It's that humans created a test where success was measured by achieving an objective, deliberately relaxed some of the normal security controls to measure the system's capabilities, and underestimated how effective the model would be at finding an unexpected path to success."
The OpenAI agent thatbroke into tech firm Hugging Facewent on a dayslong hacking spree that OpenAIdidn't notice until well after the threat was contained and the FBI was alerted, according to people familiar with the investigation. The agent -- a program capable of making decisions and executing complex tasks with little or no human oversight -- attempted to break out of its isolated testing environment at OpenAI around July 9, according to two of the people. The intrusion at Hugging Face, which operates as a repository for AI tools and models, began two days later on July 11 and lasted until July 13, said Thomas Wolf, Hugging Face's co-founder. It took several more days for OpenAI to realize its agent was behind the hack, and the two companies only communicated about it for the first time on or around July 20, according to Wolf and three of the people familiar with the investigation.
OpenAI's public disclosure, on July 21, thatone of its agents had slipped out of control and carried out the break-in at Hugging Facedrew global attention. But many details of the hack, including how long the agent went rogue and OpenAI's belated knowledge of it, are being reported here for the first time. Hugging Face is preparing a public timeline of the hack, Wolf said, adding that he could not speak to what happened at OpenAI. In a statement, OpenAI said the hac
How did an OpenAI test end up hitting Hugging Face like this?
OpenAI was testing GPT-5.6 Sol and a more powerful unreleased model using ExploitGym, a benchmark that challenges AI systems to find and exploit software vulnerabilities. The company removed some cybersecurity safeguards that would normally prevent potentially dangerous actions while relying on an isolated environment to keep the models away from the wider internet.
According to OpenAI's postmortem, the models discovered a previously unknown vulnerability in third-party software used to proxy and cache software packages. They exploited it, escalated their privileges and moved through OpenAI's research infrastructure until they reached a machine with public internet access.
Hugging Face became a target because the models identified it as a possible source of information that could help them complete the ExploitGym challenges. OpenAI said at least one attack chain involved stolen credentials and previously unknown vulnerabilities that eventually enabled the models to execute remote code on Hugging Face systems and access test solutions stored in a production database.
In their disclosure, Hugging Face representatives said the company recorded more than 17,000 actions during the intrusion, but they couldn't initially explain who or what was behind it. OpenAI's subsequent disclosure supplied that missing piece: Its models had broken out of their test environment and gone looking for the answers elsewhere.
Rather than harboring any malicious intent, the AI models simply wanted to find out more information so they could complete their task.
(Image credit: wildpixel/ Getty Images)
Did the AI really "escape"?
It's notable that the models found a flaw in the infrastructure designed to contain an AI and used it to reach the public internet. Describing the models as having "gone rogue," however, risks assigning them unsupported motivations, Buckley said.
"I think I'd be wary of jumping to "rogue AI,"" Buckley said. "The models didn't develop their own agenda or decide to attack Hugging Face while twirling their digital moustache."
Buckley compared it to asking a dog to fetch a ball while leaving the garden gate open. "If the easiest ball for it to find is in the park down the road, that's where it'll head," he said. "You wouldn't say the dog had gone rogue; you'd just say you underestimated how literally it would pursue the task."
Daniel Hulme, entrepreneur in residence at University College London and CEO of AI safety company Conscium, agreed that the models shouldn't be assigned human-like motivations. "Models don't have intent; humans have the intent, and we train models with goals in mind," he told Live Science
The capability may matter more than the motive
What matters more than the models' supposed motives is what they managed to accomplish while pursuing their assigned task.
"The genuinely significant point is that the models appear to have chained together multiple vulnerabilities across different systems and sustained a complex sequence of actions," Buckley said. "That demonstrates a level of capability that security professionals should take seriously."
The lesson isn't that AI has become malicious. Instead, it's that increasingly capable systems will exploit opportunities that humans fail to anticipate.
Oli Buckley, professor in cybersecurity at Loughborough University
Katerina Mitrokotsa, a professor of cybersecurity and applied cryptography at the University of St. Gallen in Switzerland, said the containment failure is particularly concerning because another company ultimately paid the price.
"What concerns me most is who ended up affected," Mitrokotsa told Live Science. "The victim was not the company running the test, but a third party. This is the scenario security researchers have warned about for some time: that an AI agent's escape does not necessarily stay contained to the environment in which it originated."
OpenAI representatives said they have tightened the infrastructure used for these evaluations. But Mitrokotsa warned that containment becomes harder to guarantee as models improve at performing exactly the kind of exploitation OpenAI was testing.
An AI warning — and an impressive product demonstration
There is also reason to look carefully at how the incident is being framed. OpenAI's account serves two purposes at once: It warns about the security risks posed by increasingly capable AI while demonstrating just how capable its own newest models have become.
Buckley said announcements from frontier AI companies like OpenAI or Anthropic should be viewed in the context of an industry competing to build ever-more-powerful models.
"We've seen similar high-profile capability demonstrations from Anthropic and others," he said. "That doesn't make the findings untrue, but it does mean we should separate the technical evidence from the marketing narrative."
These companies have every incentive to show both that their models are extraordinarily capable and that they are taking the risks seriously, he added. The Hugging Face incident demonstrates both that OpenAI's models carried out a complex series of operations with considerable autonomy and that its security measures failed to keep them inside the experiment.
Hulme argued that the longer-term challenge is ensuring that increasingly capable AI systems pursue their goals in ways that remain consistent with human values."Rather than seeking to control AIs, the focus should instead be on alignment," he said, adding that continuous testing will be needed to ensure systems remain aligned with their intended missions while staying secure.
The episode, the experts said, leaves OpenAI with a result that is impressive and uncomfortable in equal measure. Its models found previously unknown vulnerabilities and continued pursuing their goal well beyond the boundaries their creators expected, but none of that requires them to have developed malign intentions.
"The lesson isn't that AI has become malicious," Buckley said. "Instead, it's that increasingly capable systems will exploit opportunities that humans fail to anticipate."
In this incident, OpenAI's new models were given a hacking challenge and they were rewarded for finding a way to solve it. The humans running the experiment simply hadn't anticipated quite how far they might go.
As a Partner and Co-Founder of Predictiv and PredictivAsia, Jon specializes in management performance and organizational effectiveness for both domestic and international clients. He is an editor and author whose works include Invisible Advantage: How Intangilbles are Driving Business Performance. Learn more...
0 comments:
Post a Comment