Skip to main content
  1. Home
  2. Computing
  3. News

AI models from Anthropic and OpenAI were caught breaking the rules again

A new report states AI agents from both companies took unauthorized actions during safety tests, from hacking a website to tricking real people online.

Add as a preferred source on Google
Claude website open on laptop
Rachit Agarwal / Digital Trends

OpenAI and Anthropic have both had a rough few weeks on the AI safety front. OpenAI recently disclosed that its models broke out of a test environment and hacked into Hugging Face and four other organizations. The news prompted Anthropic to review its own testing, which revealed that Claude had also gained unauthorized access to three companies.

Now, the UK’s AI Security Institute (AISI) has disclosed a new round of incidents (via Wired). It recorded 19 unauthorized actions on the live internet across 122 test runs involving models from both companies, the most serious of which saw an agent invent fake online personas to push malicious code into a real GitHub project. OpenAI separately revealed a second incident in which one of its models hacked a real website after a third-party lab mistakenly gave it live internet access.

17 incidents tied to Anthropic’s Mythos 5

AISI traced 17 of the 19 unauthorized actions to Anthropic’s Mythos 5 model, with the remaining two tied to OpenAI’s GPT 5.6 Sol. The GitHub incident was one of the 17, and it didn’t end when a human reviewer rejected the submission. The agent posted a summary of its progress publicly, inviting other automated systems to pick up where it left off, an attempt at what AISI calls prompt injection. A separate agent later found that message, used it, and continued the work.

On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations.

The behaviour came mostly from one model (Anthropic’s Mythos 5), with a small number of events from… pic.twitter.com/SPnA4Ekkwq

— AI Security Institute (AISI) (@AISecurityInst) August 4, 2026

AISI says it deliberately gave the models internet access and relaxed some safety protections to test their capabilities, but never instructed the agents to target real people or organizations. The institute says it’s still unclear whether the agents understood they’d gone beyond the scope of the simulation.

Another accidental breach at OpenAI

A second incident, disclosed by OpenAI the same day, started with a mistake at Irregular, a third-party lab OpenAI hired to run its cybersecurity tests. Irregular meant to keep its evaluation model confined to an isolated sandbox, but a configuration error gave the model direct access to the live internet. Once out, it exploited a vulnerability to break into a real website, then found and used credentials to operate the site it had just hacked. OpenAI hasn’t named the website or detailed what the model did with its access.

Recommended Videos

Both companies say the new incidents happened under deliberately loosened conditions that don’t reflect how their public models behave. Be that as it may, that doesn’t change the fact that AI agents from two of the industry’s most closely watched companies have now slipped past their intended limits in three separate incidents within a matter of weeks. And that doesn’t bode well for an industry racing to hand AI agents more real-world tasks before proving it can keep them in check.

Pranob Mehrotra
Pranob is a seasoned tech journalist with over eight years of experience covering consumer technology. His work has been…
Study finds readers rate AI-written stories higher, but still trust the “human” label more
Researchers also found that readers could barely differentiate between AI-generated and human-written stories, though those familiar with AI tools guessed somewhat better.
updated book and AI photo

If you think AI writing is easy to spot, a new study suggests you may be overestimating your ability to tell the difference. Researchers from Villanova University have found that readers struggled to tell AI-generated stories from human-written ones and frequently rated the AI versions higher.

The study, led by Dr. Deena Skolnick Weisberg and published in the journal Judgment and Decision Making (via The Guardian), asked more than 1,600 participants to rate one of six short stories, three written by a human and the rest generated by ChatGPT, on quality and engagement. Participants were told who wrote their story, though that label wasn't always accurate.

Read more
Motherboards may be the next PC component to get a massive price increase
After soaring RAM and storage prices, motherboards could be the next PC component to get significantly more expensive, according to a new supply chain report.
Gigabyte X870E AORUS INFINITY NEXT motherboard on display

Building a PC from scratch has already become a lot more expensive, thanks to skyrocketing RAM and storage prices driven by rising AI data center demand. Now, a new report suggests motherboard prices could soon push costs even higher, with boards from Asus, MSI, and Gigabyte rumored to see price hikes of at least 50 percent.

Your next PC upgrade could get a lot pricier

Read more
Deepfake bosses are crashing video calls, and researchers are trying to expose them
Seeing your boss on video no longer proves they are real
Researchers at Fraunhofer SIT are fighting against deepfake meeting video calls

Entering into a meeting with your boss and several familiar coworkers inside a video conference is the next area vulnerable to cybercrimes, and it's all because of deepfake technology. Researchers at the Fraunhofer Institute for Secure Information Technology SIT are working on a real-time warning system designed to identify attempted fraud during corporate video conferences.

The report states that criminals are increasingly targeting video meetings for identity theft and financial scams, taking advantage of the trust people place in familiar faces and voices. The intention is to flag suspicious activity while the meeting is still taking place. This gives an employee a chance to stop before following an expensive instruction from an AI-generated executive.

Read more