Anthropic says Claude hacked real companies during AI safety tests



In one incident, Claude Opus 4.7 hacked into an outside production database over the internet, and continued the hack even after realizing the company it was attacking was real. 

In another occurrence, Claude Mythos 5 uploaded a bogus Python package to PyPI, the public Python repository. The malicious package was downloaded and installed by 15 real-world companies, including a security firm, Anthropic admitted.

In the third attack, an internal Claude model that was never released used “basic and well-known cyberattack techniques” to hack a company’s “internet-facing application,” assuming it was part of the “capture-the-flag” exercise. The silver lining is that the Claude model stopped attacking once it realized the target company was real.

In each case, the Claude models were supposed to be operating in walled-off test environments with no internet access. But Anthropic now says the models actually could reach the internet due to a human “misconfiguration,” leading the models to believe that the real companies they were attacking were part of their training exercises.

So, are we talking another case of “frontier” AI models run amok? For its part, Anthropic is blaming human error for the real-world hack attacks, not the models themselves.

“We saw no evidence in any run described here of a model pursuing a goal of its own,” the Anthropic post-mortem said. “Instead, the models did what their evaluation asked — though in most cases, they did so while holding a false belief about whether the environment was real.”



Source link