Anthropic’s Claude hacked three real-life companies during security capabilities test — test environment with internet access and unwitting targets’ lax cyberse

Anthropic's Claude hacked three real-life companies during security capabilities test — test environment with internet access and unwitting targets' lax cyberse

In a twist of irony, one such system belonged to a security vendor that scans PyPI for malware, and lo and behold, promptly failed to find Claude's booby-trap and ran it. Once Claude presumably had remote code execution privileges, it used the credentials it found for further infiltration. The amusing bit is that Claude had no idea this company existed and didn't target it; the downloads just happened because the package was up and live for a short while.

According to Anthropic, the bot did detect it was acting on the real internet and even said that publishing a package like this was "NOT okay." However, it talked itself into believing it was in a test environment as it didn't recognize the real SSL certificates for the connections. It even believed the 2026 calendar date on the systems "proved" the environment was staged. It even recognized the systems that installed the malware as part of the experiment.

As for the third incident, Anthropic isn't saying much, other than Claude scanned 9,000 real live potential alternative targets once it noticed the intended one wasn't reachable. One of them reportedly had a live page with debugging information and was vulnerable to plain ol' SQL injection. Interestingly, this time around, once Claude noticed that the servers it was accessing resided on a cloud environment and not on the local network, it stopped the attack.

For its part, Anthropic recognizes that despite Claude following the instructions for the objectives, the fact that it stopped by itself once it found the target was real in only one of the cases is food for some thought. The firm says it's talking to METR for a third-party review, and it needs to "better co-design evaluation environments."

Oddly enough, Anthropic believes Claude probably wouldn't have gone online "if the prompt had clearly explained which systems were in and out of scope for the evaluation," while also stating the incidents were "closer to a harness and operational failure than a model alignment failure."

Follow Tom's Hardware on Google News , or add us as a preferred source , to get our latest news, analysis, & reviews in your feeds.

Bruno Ferreira is a contributing writer for Tom's Hardware. He has decades of experience with PC hardware and assorted sundries, alongside a career as a developer. He's obsessed with detail and has a tendency to ramble on the topics he loves. When not doing that, he's usually playing games, or at live music shows and festivals. ","collapsible":{"enabled":true,"maxHeight":250,"readMoreText":"Read more","readLessText":"Read less"}}), "https://slice.vanilla.futurecdn.net/13-4-25/js/authorBio.js"); } else { console.error('%c FTE ','background: #9306F9; color: #ffffff','no lazy slice hydration function available'); } Bruno Ferreira Social Links Navigation Contributor Bruno Ferreira is a contributing writer for Tom's Hardware. He has decades of experience with PC hardware and assorted sundries, alongside a career as a developer. He's obsessed with detail and has a tendency to ramble on the topics he loves. When not doing that, he's usually playing games, or at live music shows and festivals.

alrighty_then Machine speed hacking is here. The bad guys will have a field-day because there is just no way that all the companies & nation states will get their act together and patch/scan/fix with the fronteir models BEFORE the blackhats have started running their own automated attacks with their own comparable models. It will be a ransomware nightmare. Maybe shutting down bitcoin would be eaiser than trying to stop the hacks…seriously! Reply

warezme What I'm reading is not definitive instructional control values provided to the model and an OPEN network with guardrails removed. Sounds like an intentional release of the model to garner both funding support and headlines. Reply

psyconz alrighty_then said: Machine speed hacking is here. The bad guys will have a field-day because there is just no way that all the companies & nation states will get their act together and patch/scan/fix with the fronteir models BEFORE the blackhats have started running their own automated attacks with their own comparable models. It will be a ransomware nightmare. Maybe shutting down bitcoin would be eaiser than trying to stop the hacks…seriously! That's a fair take. Just curious how anyone is supposed to shut down a decentralized system like Bitcoin. The only thing that could happen is that the bottom drops out of the value of it and people stop using it as much, but bitcoin is here to stay. Reply

TommyCC It shows closed frontier models are not 100% secure. Reply

Key considerations

  • Investor positioning can change fast
  • Volatility remains possible near catalysts
  • Macro rates and liquidity can dominate flows

Reference reading

More on this site

Informational only. No financial advice. Do your own research.

Leave a Comment