Meta’s AI model follows rivals in revealing hacks of outside systems


Weeks before OpenAI's agents hacked Hugging Face, the agents worked together to find and exploit a vulnerability in the infrastructure supporting the company's cybersecurity testing, OpenAI researchers said Wednesday.Why it matters: The new findings raise questions about how frontier AI labs are monitoring their testing environments — and the challenges safety testers are finding as they try to rein in increasingly powerful AI.Driving the news: OpenAI's internal research model, one of the models involved in the Hugging Face breach, first discovered and exploited a vulnerability in Artifactory, a third-party file repository connected to the company's testing sandbox, on May 26, two researchers said at the Black Hat cybersecurity conference on Wednesday.The genius-level actions observed in testing showed exactly what could happen if these models make it to the hands of hackers."We believe this is a watershed moment for computer security as an industry, as well as for OpenAI as a whole," Michael Dalton, a member of OpenAI's technical staff, said during the presentation.What happened: OpenAI started testing its internal research model, which is not intended for public release, on May 7.Within a day, after getting stuck while trying to complete a cybersecurity evaluation, the model reasoned that it could reach the internet indirectly. The model then discovered it could write its own files in Artifactory's shared package repository.It then left a note to other agents about its discovery.
Reported by 3 outlets — Axios, Al Jazeera, The Verge. See all sources ↓
Weeks before OpenAI's agents hacked Hugging Face, the agents worked together to find and exploit a vulnerability in the infrastructure supporting the company's cybersecurity testing, OpenAI researchers said Wednesday.Why it matters: The new findings raise questions about how frontier AI labs are monitoring their testing environments — and the challenges safety testers are finding as they try to rein in increasingly powerful AI.Driving the news: OpenAI's internal research model, one of the models involved in the Hugging Face breach, first discovered and exploited a vulnerability in Artifactory, a third-party file repository connected to the company's testing sandbox, on May 26, two researchers said at the Black Hat cybersecurity conference on Wednesday.The genius-level actions observed in testing showed exactly what could happen if these models make it to the hands of hackers."We believe this is a watershed moment for computer security as an industry, as well as for OpenAI as a whole," Michael Dalton, a member of OpenAI's technical staff, said during the presentation.What happened: OpenAI started testing its internal research model, which is not intended for public release, on May 7.Within a day, after getting stuck while trying to complete a cybersecurity evaluation, the model reasoned that it could reach the internet indirectly. The model then discovered it could write its own files in Artifactory's shared package repository.It then left a note to other agents about its discovery. Soon, more agents started leaving notes for each other in the repository, creating a de facto message board where the agents collaborated and traded information about their findings, including new vulnerabilities they found. Zoom in: The agents uncovered a variety of vulnerabilities in Artifactory, including a remote code execution flaw and another that gave them administrator privileges.
Read the full report at Axios ↗
Why it matters
3 outlets are covering this world story — one to watch as reporting develops.
- What's the story?
- Weeks before OpenAI's agents hacked Hugging Face, the agents worked together to find and exploit a vulnerability in the infrastructure supporting the company's cybersecurity testing, OpenAI researchers said Wednesday.Why it matters: The new findings raise questions about how frontier AI labs are monitoring their testing environments — and the challenges safety testers are finding as they try to rein in increasingly powerful AI.Driving the news: OpenAI's internal research model, one of the models involved in the Hugging Face breach, first discovered and exploited a vulnerability in Artifactory, a third-party file repository connected to the company's testing sandbox, on May 26, two researchers said at the Black Hat cybersecurity conference on Wednesday.The genius-level actions observed in testing showed exactly what could happen if these models make it to the hands of hackers."We believe this is a watershed moment for computer security as an industry, as well as for OpenAI as a whole," Michael Dalton, a member of OpenAI's technical staff, said during the presentation.What happened: OpenAI started testing its internal research model, which is not intended for public release, on May 7.Within a day, after getting stuck while trying to complete a cybersecurity evaluation, the model reasoned that it could reach the internet indirectly. The model then discovered it could write its own files in Artifactory's shared package repository.It then left a note to other agents about its discovery.
- How widely is it covered?
- 3 outlets, average source rating 6.7/10.
- When was it last updated?
- 3m ago.
How outlets are framing the same story
Here's how each outlet is covering the story — compare their headlines and timing at a glance.
- Coverage card1 outlet1CoverageScouting report
How OpenAI's agents broke out of testing to hack Hugging Face
Sources1TypeCoverageAxios
- Coverage card1 outlet2CoverageScouting report
Meta’s AI model follows rivals in revealing hacks of outside systems
Sources1TypeCoverageAl Jazeera
- Coverage card1 outlet3CoverageScouting report
Rogue AI agents created fake online identities in another hacking attempt
Sources1TypeCoverageThe Verge