
Regarding the huge reported number of incidents, Axios' sources point out that Anthropic and other labs conduct hundreds of thousands of test runs on their models; therefore, even a small percentage of misaligned behavior can add up to tens of thousands of incidents. Anthropic, for example, combed through 141,006 evaluation runs in which Claude had internet access and found three incidents in which Claude hacked three real-life companies during security capabilities testing .
Anthropic said it considers those incidents closer to a harness and operational failure than a model alignment failure, as the models were told they had no internet access while in fact being misconfigured to have it. Outside of Anthropic and OpenAI, other AI models have their share of incidents. Google confirmed a report that its Gemini models hacked three companies earlier this year.
Some experts Axios interviewed believe that these incidents are one-off, and expect future disclosures to be less severe thanks to improved controls. They also say there are simple fixes that would help labs avoid parts of what made the episodes look so dangerous to outsiders. On the other hand, other stakeholders expressed limited confidence that AI companies can prevent all problematic model behavior, saying that this would require the impossible task of anticipating every possible way the models might go off track. Either way, experts believe some misaligned behavior is expected as labs test new models, and bringing that risk to zero may not be feasible.
These incidents have intensified calls for guardrails across the AI industry. In an essay backed by OpenAI, Google DeepMind, and Microsoft executives, Anthropic CEO Dario Amodei urged Washington to pace AI development over fears that agents could spiral out of control, warning of a potential AI-powered botnet swarm that could take over the entire internet . President Donald Trump has repeatedly dismissed calls for regulations as hoaxes and has announced plans to establish an “AI force” to cherish, help, and watch over the AI industry as it grows.
Follow Tom's Hardware on Google News , or add us as a preferred source , to get our latest news, analysis, & reviews in your feeds.
Etiido Uko Social Links Navigation News Contributor Etiido Uko is a news contributor for Tom's Hardware covering the latest updates in big tech and the PC industry. He is a mechanical engineer and senior technical writer with over nine years of experience in documentation and reporting. He is deeply passionate about all things engineering and technology, and is an expert in gadgets, manufacturing, robotics, automotive, and aerospace.
Key considerations
- Investor positioning can change fast
- Volatility remains possible near catalysts
- Macro rates and liquidity can dominate flows
Reference reading
- https://www.tomshardware.com/tech-industry/artificial-intelligence/SPONSORED_LINK_URL
- https://www.tomshardware.com/tech-industry/artificial-intelligence/openai-and-anthropic-are-reportedly-investigating-tens-of-thousands-of-ai-security-incidents-openai-pauses-testing-after-ai-kill-switch-fails-to-stop-a-rogue-agent-report-says-problem-is-orders-of-magnitude-more-complex-than-what-is-publicly-known#main
- https://www.tomshardware.com/membership
- From Enablement to Execution, Egypt’s AI Ecosystem Reaches Production Scale
- Nvidia DLSS 5 power draw hits 647W as power connector runs hotter than the GPU die, upscales frame rates and flame temps on 16-pin
- How Open Science Can Help Researchers Prepare for the Next Pandemic
- North Korea named as primary suspect in $387 million Bitget crypto hack
- G.Skill wins PC enthusiast as 'customer for life' by simply honoring its warranty replacement policy
Informational only. No financial advice. Do your own research.