How OpenAI's agents broke out of testing to hack Hugging Face

Weeks before OpenAI's agents hacked Hugging Face, the agents worked together to find and exploit a vulnerability in the infrastructure supporting the company's cybersecurity testing, OpenAI researchers said Wednesday.Why it matters: The new findings raise questions about how frontier AI labs are monitoring their testing environments — and the challenges safety testers are finding as they try to rein in increasingly powerful AI.Driving the news: OpenAI's internal research model, one of the models involved in the Hugging Face breach, first discovered and exploited a vulnerability in Artifactory, a third-party file repository connected to the company's testing sandbox, on May 26, two researchers said at the Black Hat cybersecurity conference on Wednesday.The genius-level actions observed in testing showed exactly what could happen if these models make it to the hands of hackers."We believe this is a watershed moment for computer security as an industry, as well as for OpenAI as a whole," Michael Dalton, a member of OpenAI's technical staff, said during the presentation.What happened: OpenAI started testing its internal research model, which is not intended for public release, on May 7.Within a day, after getting stuck while trying to complete a cybersecurity evaluation, the model reasoned that it could reach the internet indirectly. The model then discovered it could write its own files in Artifactory's shared package repository.It then left a note to other agents about its discovery.
Reported by 1 outlet — Axios. See all sources ↓
Weeks before OpenAI's agents hacked Hugging Face, the agents worked together to find and exploit a vulnerability in the infrastructure supporting the company's cybersecurity testing, OpenAI researchers said Wednesday.Why it matters: The new findings raise questions about how frontier AI labs are monitoring their testing environments — and the challenges safety testers are finding as they try to rein in increasingly powerful AI.Driving the news: OpenAI's internal research model, one of the models involved in the Hugging Face breach, first discovered and exploited a vulnerability in Artifactory, a third-party file repository connected to the company's testing sandbox, on May 26, two researchers said at the Black Hat cybersecurity conference on Wednesday.The genius-level actions observed in testing showed exactly what could happen if these models make it to the hands of hackers."We believe this is a watershed moment for computer security as an industry, as well as for OpenAI as a whole," Michael Dalton, a member of OpenAI's technical staff, said during the presentation.What happened: OpenAI started testing its internal research model, which is not intended for public release, on May 7.Within a day, after getting stuck while trying to complete a cybersecurity evaluation, the model reasoned that it could reach the internet indirectly. The model then discovered it could write its own files in Artifactory's shared package repository.It then left a note to other agents about its discovery. Soon, more agents started leaving notes for each other in the repository, creating a de facto message board where the agents collaborated and traded information about their findings, including new vulnerabilities they found. Zoom in: The agents uncovered a variety of vulnerabilities in Artifactory, including a remote code execution flaw and another that gave them administrator privileges.
Read the full report at Axios ↗
Why it matters
A world story we're tracking; its significance and source trust firm up as more outlets confirm it.
- What's the story?
- Weeks before OpenAI's agents hacked Hugging Face, the agents worked together to find and exploit a vulnerability in the infrastructure supporting the company's cybersecurity testing, OpenAI researchers said Wednesday.Why it matters: The new findings raise questions about how frontier AI labs are monitoring their testing environments — and the challenges safety testers are finding as they try to rein in increasingly powerful AI.Driving the news: OpenAI's internal research model, one of the models involved in the Hugging Face breach, first discovered and exploited a vulnerability in Artifactory, a third-party file repository connected to the company's testing sandbox, on May 26, two researchers said at the Black Hat cybersecurity conference on Wednesday.The genius-level actions observed in testing showed exactly what could happen if these models make it to the hands of hackers."We believe this is a watershed moment for computer security as an industry, as well as for OpenAI as a whole," Michael Dalton, a member of OpenAI's technical staff, said during the presentation.What happened: OpenAI started testing its internal research model, which is not intended for public release, on May 7.Within a day, after getting stuck while trying to complete a cybersecurity evaluation, the model reasoned that it could reach the internet indirectly. The model then discovered it could write its own files in Artifactory's shared package repository.It then left a note to other agents about its discovery.
- How widely is it covered?
- 1 outlet, average source rating 7.0/10.
- When was it last updated?
- 12m ago.
How outlets are framing the same story
Here's how each outlet is covering the story — compare their headlines and timing at a glance.
- Coverage card1 outlet1CoverageScouting report
How OpenAI's agents broke out of testing to hack Hugging Face
Sources1TypeCoverageAxios