OpenAI确认其预发布AI模型在内部测试中攻击Hugging Face

原文归档。生成时间:2026-07-21 22:46 UTC

返回当日简报

OpenAI says it accidentally hacked Hugging Face with a new AI system

The Verge 原始链接

OpenAI says its AI models mistakenly breached open-source AI platform Hugging Face during internal testing. In a blog post on Tuesday, OpenAI writes that GPT-5.6 Sol and "an even more capable pre-release model" discovered vulnerabilities within their sandboxed testing environment, allowing them to gain access to the internet and target Hugging Face. On July 16th, […]

OpenAI says its AI models mistakenly breached open-source AI platform Hugging Face during internal testing. In a blog post on Tuesday, OpenAI writes that GPT-5. 6 Sol and “an even more capable pre-release model” discovered vulnerabilities within their sandboxed testing environment, allowing them to gain access to the internet and target Hugging Face.

OpenAI says it accidentally hacked Hugging Face with a new AI system The announcement about a serious security issue oddly reads like an advertisement for how capable OpenAI’s technology is. The announcement about a serious security issue oddly reads like an advertisement for how capable OpenAI’s technology is.

On July 16th, Hugging Face disclosed a security incident that it says was driven by “an autonomous AI agent system. ” Hugging Face’s AI agents detected and stopped the breach, which OpenAI has now admitted occurred during an evaluation of its models’ cybersecurity capabilities.

OpenAI says “all evidence suggests that the models were hyperfocused on finding a solution for ExploitGym,” a benchmark system that measures whether AI models can turn security vulnerabilities into exploits. As part of efforts to complete the evaluation, the AI models gained access to the internet by exploiting a zero-day vulnerability in the sandboxed environment.

From there, OpenAI says its models “inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym,” and then “searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation:” In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers.

But as serious as this incident is, OpenAI appears to be using the “unprecedented” attack as an opportunity to make its AI systems look good — especially as it competes with cybersecurity rivals, like Anthropic’s Mythos and Gemini Flash 3. 5 Cyber. OpenAI’s blog post has a chart showing how GPT-5.

6 Sol is getting better at sustaining multi-step cyber operations, and also encourages enterprise customers to sign up to access its “Cyber” security model. OpenAI adds that it’s now working with Hugging Face to investigate the security incident, and will implement new controls within its research environment.

Most Popular - The FCC is planning to retroactively ban disguised DJI gadgets - The Light Flip is a minimalist flip phone with a point to prove - China delivers a one-two punch to America’s AI dominance - Garmin’s new screen-free fitness tracker doesn’t require a subscription - Apple’s rumored ‘Upgrade’ program brings lease-to-own pricing for iPhones, Macs, and iPads

OpenAI says Hugging Face was breached by its own pre-release models

TechCrunch AI 原始链接

OpenAI has come forward to claim responsibility for the Hugging Face breach, saying it was the result of internal testing gone awry.

OpenAI admitted Tuesday that one of its AI models breached the systems of Hugging Face, the unaffiliated AI hosting platform, during an internal cybersecurity test that went awry. The models reportedly escaped their isolated testing environment and reached Hugging Face’s systems from there.

Hugging Face initially attributed the breach to an “external AI agent. ” In a blog post published Tuesday afternoon, OpenAI detailed the steps that led the models to compromise the service. “After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.

6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities,” the post reads. In particular, the breach appears to have focused on ExploitGym, a publicly hosted benchmark measuring models’ ability to execute attacks based on existing vulnerabilities.

Benchmarks like ExploitGym are commonly used in model training to refine specific skills, but this is the first known incident in which that testing resulted in an actual cyberattack. In this case, the model in question should not have even had internet access, outside of a specific tool that enabled models to install software packages they might need to complete their task.

Instead, the model was able to find an undisclosed vulnerability in the package-installer program, which it used to access the broader internet at will. “The models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal,” OpenAI’s post reads.

“After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.

” Ultimately, the models found vulnerabilities in Hugging Face’s infrastructure that allowed them to “obtain test solutions directly from Hugging Face’s production database,” effectively providing the answers to the benchmark. For Hugging Face, the apparent result was a sophisticated and aggressive cyberattack, with “many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services,” as the company stated in its initial disclosure.

OpenAI has identified and reported the vulnerabilities in the package installer and is working with Hugging Face to investigate the incident further. The company also said it would implement new controls on both model testing and the related infrastructure, meant to prevent similar incidents in the future.

It’s unclear whether OpenAI will face any legal consequences as a result of the breach, although it’s likely that the models’ actions violated the Computer Fraude and Abuse Act. Nevertheless, the result is an unusually vivid illustration of the power and dangers of frontier AI models operating on long time horizons.

As OpenAI researcher Micah Carroll posted in response to the news, “If this doesn’t convince you that misalignment risks are going to be a key concern going forward, I don’t know what will. ”

OpenAI and Hugging Face partner to address security incident during model evaluation

OpenAI Blog 原始链接

OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities and lessons for defenders.

OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities and lessons for defenders.