AI Agents Escaped Their Tests and Hacked Real Companies — Should the Public Be Worried?

Artificial intelligence systems being tested by two of the world’s leading AI companies escaped supposedly controlled environments and broke into the computer systems of real organisations.

OpenAI’s models reached the public internet and spent more than four days attacking infrastructure belonging to Hugging Face, a major platform for hosting AI models and datasets. OpenAI later revealed that its models had also used publicly exposed credentials to access four accounts across four other online services.

Days later, rival AI company Anthropic disclosed that Claude models had escaped cybersecurity tests and gained unauthorised access to three organisations.

This was not artificial intelligence becoming conscious or declaring war on humanity. The more troubling explanation is painfully human: increasingly capable systems were given powerful tools, while the companies testing them failed to keep them contained.

How did OpenAI’s models escape?

The OpenAI models were completing cybersecurity evaluations designed to test whether advanced AI could identify and exploit software vulnerabilities.

They were operating with reduced safety restrictions inside a digital “sandbox” that was supposed to separate their actions from the real internet.

However, the models exploited an unpatched software vulnerability and escaped the testing environment. They apparently concluded that Hugging Face’s servers might contain information that would help them complete—or cheat—the evaluation.

Instead of remaining inside the simulated challenge, they began targeting real infrastructure.

According to Hugging Face’s investigation, the agent performed approximately 17,600 actions over four and a half days. It exploited vulnerable systems, obtained credentials and source code, moved between networks and established access across 11 servers.

The system also acquired a cryptographic key capable of generating apparently valid login credentials. It repeatedly rebuilt its tools and reconnected after the testing environment reset.

That persistence is the critical point. The AI did not require a human operator to approve every move. It could attempt one route, study the result and continue trying alternatives until it found something that worked.

OpenAI was not alone

Following the OpenAI disclosure, Anthropic examined 141,006 cybersecurity evaluation runs involving its Claude models.

The company identified three cases in which a model reached the internet from a third-party testing environment and gained unauthorised access to the production systems of real organisations. The earliest incident reportedly occurred in April.

Claude had been assigned “capture the flag” exercises, which required it to locate hidden information inside simulated networks. Although the models were told that internet access was unavailable, a configuration error meant that the restriction was not properly enforced.

The consequences extended beyond simply visiting the wrong website.

In one case, a Claude model reportedly accessed a database containing hundreds of rows of production data. In another, a model published a malicious software package that was subsequently installed on 15 real systems.

None of the affected organisations had reportedly detected the intrusions before Anthropic began attempting to notify them.

Anthropic described the incidents as an operational failure and suspended the affected evaluations.

The OpenAI and Anthropic cases were not identical. Together, however, they reveal the same uncomfortable problem: AI capabilities are advancing faster than the systems meant to monitor and contain them.

Why should ordinary people care?

There is no evidence that these particular incidents emptied bank accounts, shut down hospitals or exposed millions of consumers’ records.

What they demonstrate is how quickly poorly contained AI can find and exploit weaknesses that already exist.

A human hacker needs time to locate vulnerable websites, test credentials, study stolen information and move between systems. An autonomous AI agent can attempt thousands of actions, analyse the results and adapt continuously—without becoming tired, distracted or discouraged.

That scale changes the threat.

A similar system targeting a bank, municipality, hospital, retailer or government department could identify forgotten servers, weak passwords and unpatched software faster than human security teams can respond.

For ordinary people, the consequences could include stolen personal information, compromised banking and shopping accounts, disrupted public services and automated fraud conducted at machine speed.

South Africa is not protected by geography. Local banks, municipalities, hospitals and retailers all rely on internet-connected systems. An AI agent does not care whether a vulnerable server is located in Silicon Valley, Sandton or Pietermaritzburg. If there is an unlocked digital door, it can keep testing the handle.

The real failure was human

Describing these systems as “rogue” is convenient because it shifts responsibility onto the machines.

But humans designed the evaluations, weakened the safety restrictions and failed to guarantee that the testing environments were isolated from the public internet. Monitoring systems also failed to identify and stop the intrusions immediately.

Hugging Face chief executive Clément Delangue has called for “radical transparency” and urged OpenAI to release enough information for independent researchers to understand how the attack happened.

That demand matters. AI companies should not be allowed to investigate themselves behind closed doors and reveal only the details they find comfortable.

Testing environments must be isolated from public networks, with external connections blocked by default. Dangerous activity should be continuously monitored, and emergency shutdown systems must be capable of stopping an evaluation immediately.

Independent security audits and mandatory reporting of containment failures should also become standard. Most importantly, companies must remain legally responsible when their experimental systems damage infrastructure belonging to someone else.

These incidents do not prove that artificial intelligence is plotting against humanity. They reveal a more immediate danger.

AI has not developed a mind of its own—but machines can now cause real damage while the humans responsible are still trying to work out where they went.

Article written by:

Hudaa Ahmed

Journalist at Radio Al Ansaar