U.K. government reports OpenAI, Anthropic models attempted to hack companies

Two third-party testing firms said Tuesday that they've uncovered more instances where Anthropic and OpenAI's most advanced models tried — and sometimes succeeded in — compromising third-party systems last month. Why it matters: The incidents add to a growing string of disclosures showing frontier AI models taking unsanctioned actions against real people, organizations and online services while trying to complete cybersecurity evaluations.State of play: The U.K.
Reported by 1 outlet — Axios. See all sources ↓
Two third-party testing firms said Tuesday that they've uncovered more instances where Anthropic and OpenAI's most advanced models tried — and sometimes succeeded in — compromising third-party systems last month. Why it matters: The incidents add to a growing string of disclosures showing frontier AI models taking unsanctioned actions against real people, organizations and online services while trying to complete cybersecurity evaluations.State of play: The U.K. AI Security Institute, a government body that conducts safety and security testing of top AI models, said Tuesday, that it documented 19 instances of Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol trying to hack people and companies during safety testing last month.Mythos drove 17 of those actions while GPT-5.6 Sol was behind the other two.The Institute says the models accessed GitHub during testing and created fake GitHub identities, socially engineered maintainers, planted prompt injections and sent deceptive emails.GitHub has confirmed that this violated their terms of service.GitHub and the Security Institute worked together to remove artifacts left behind by the agent, and to notify the GitHub users the model interacted with.OpenAI also said in a blog post Tuesday that its third-party safety partner, Irregular, uncovered a case where its models were mistakenly given access to the internet and broke into a real website that had the same name as the fictional company in the simulated environment. OpenAI's Irregular incident follows Anthropic's incident, shared last week.
Read the full report at Axios ↗
Why it matters
A world story we're tracking; its significance and source trust firm up as more outlets confirm it.
- What's the story?
- Two third-party testing firms said Tuesday that they've uncovered more instances where Anthropic and OpenAI's most advanced models tried — and sometimes succeeded in — compromising third-party systems last month. Why it matters: The incidents add to a growing string of disclosures showing frontier AI models taking unsanctioned actions against real people, organizations and online services while trying to complete cybersecurity evaluations.State of play: The U.K.
- How widely is it covered?
- 1 outlet, average source rating 7.0/10.
- When was it last updated?
- 7m ago.
How outlets are framing the same story
Here's how each outlet is covering the story — compare their headlines and timing at a glance.
- Coverage card1 outlet1CoverageScouting report
U.K. government reports OpenAI, Anthropic models attempted to hack companies
Sources1TypeCoverageAxios