Big mistake in AI Cybersecurity Evals: When Anthropic Cloud AI inadvertently attacked real-world systems

Anthropic logo featured image for cybersecurity evaluations incident news post

A shocking news has emerged regarding the security of artificial intelligence. A leading AI lab Anthropic recently said in its statement that a major glitch occurred during their cyber-security testing, due to which Claude AI The models unknowingly gained unauthorized access to the systems of three different real companies/organizations using the Internet.

This revelation came after Anthropic conducted an in-depth review of over 141,000 cybersecurity evaluation runs following a security incident recently reported by OpenAI. Let’s try to understand what happened, how dangerous it was, and what it means for the AI ​​industry.

How did this accident happen?

Anthropic runs simulation exercises like “Capture-The-Flag” (CTF) to test the cyber capabilities of its AI models, such as Claude Opus 4.7 and Mythos 5. In this exercise, the AI ​​model is tasked with navigating a fictitious company’s systems and finding hidden “flags” or data.

The model was clearly informed in the prompt that it was in an isolated simulation and did not have internet access. However, third-party evaluation partners Irregular Due to a misunderstanding of the network configuration with the AI ​​model, it actually had live internet access.

Because the model believed it was operating in a simulation, it also mistook live web platforms on the real Internet as part of its game/task and attacked them.

Three major events

According to Anthropic’s audit report, a total of three separate incidents have come to light. Claude Opus 4.7 In one simulation, the name of a fictitious company was identical to a real-world domain. Claude Opus 4.7 exploited infrastructure flaws in the real company, such as weak passwords, and obtained several hundred rows of data and credentials from their production database.

During another evaluation, Claude Mythos 5 read a fake developer guide. He went to the Python Package Index (PyPI) and created a genuine package and inserted malware-like code into it. This package remained live for approximately one hour and was downloaded to 15 systems. When a cybersecurity company’s scanner scanned it, the AI’s code extracted the company’s credentials and sent them to the AI.

In the third incident, an internal research model scanned nearly 9,000 web targets and hacked a live cloud application using basic techniques like SQL injection. But the model’s most remarkable feat was that when it realized the target was real and not part of a simulation, it responded.Stopped the attack on his own।

Read More:

Had Claude become a rebel?

Anthropic clarified that this was not a case of the AI ​​jailbreaking itself. The AI ​​had not set any specific goals of its own. It was simply doing what it was told to do in the task it was given.

The problem is not in the fidelity of the model, but Failure of the system’s operational setup and internet isolationThe model was misinformed that internet access was down, causing it to attack real-world systems as fake.

Anthropic’s tough steps and lessons

Upon learning of this incident, Anthropic immediately paused all cyber evaluation runs. Affected companies and partners have been notified and are assisting with data security.

For the future, Anthropic is taking the following steps: Military-grade network security for the testing environment and real-time log monitoring will be provided. Free agency METR  third-party review of the entire matter is being conducted in collaboration with. Models will be explicitly told which IP addresses and domains are in scope and which are out.

This incident proves that as AI agents become more powerful and autonomous, their testing protocols must also become more stringent. Applying AI to cyber evaluations without human oversight or robust network isolation can pose real-world risks. In this new era of AI security, not only model training but also the security of the testing environment has become a top priority.

Leave a Comment

Your email address will not be published. Required fields are marked *