
Regarding the huge reported number of incidents, Axios' sources point out that Anthropic and other labs conduct hundreds of thousands of test runs on their models; therefore, even a small percentage of misaligned behavior can add up to tens of thousands of incidents. Anthropic, for example, combed through 141,006 evaluation runs in which Claude had internet access and found three incidents in which Claude hacked three real-life companies during security capabilities testing .
Anthropic said it considers those incidents closer to a harness and operational failure than a model alignment failure, as the models were told they had no internet access while in fact being misconfigured to have it. Outside of Anthropic and OpenAI, other AI models have their share of incidents. Google confirmed a report that its Gemini models hacked three companies earlier this year.
Some experts Axios interviewed believe that these incidents are one-off, and expect future disclosures to be less severe thanks to improved controls. They also say there are simple fixes that would help labs avoid parts of what made the episodes look so dangerous to outsiders. On the other hand, other stakeholders expressed limited confidence that AI companies can prevent all problematic model behavior, saying that this would require the impossible task of anticipating every possible way the models might go off track. Either way, experts believe some misaligned behavior is expected as labs test new models, and bringing that risk to zero may not be feasible.
These incidents have intensified calls for guardrails across the AI industry. In an essay backed by OpenAI, Google DeepMind, and Microsoft executives, Anthropic CEO Dario Amodei urged Washington to pace AI development over fears that agents could spiral out of control, warning of a potential AI-powered botnet swarm that could take over the entire internet . President Donald Trump has repeatedly dismissed calls for regulations as hoaxes and has announced plans to establish an “AI force” to cherish, help, and watch over the AI industry as it grows.
Follow Tom's Hardware on Google News , or add us as a preferred source , to get our latest news, analysis, & reviews in your feeds.
Etiido Uko Social Links Navigation News Contributor Etiido Uko is a news contributor for Tom's Hardware covering the latest updates in big tech and the PC industry. He is a mechanical engineer and senior technical writer with over nine years of experience in documentation and reporting. He is deeply passionate about all things engineering and technology, and is an expert in gadgets, manufacturing, robotics, automotive, and aerospace.
hotaru251 Technology has always had the issue of advancing w/o guardrails in place. "ai" is just the extreme case in its faster and better than humans at finding workarounds/cracks. Reply
me-three hotaru251 said: Technology has always had the issue of advancing w/o guardrails in place. One reason I don't worry (as much) about "doomsday" AI is because all eyes are on it. We're all hyped about it. If duper-AI ever takes a dump, we're all ready to reach for headcovers fast. OH NOES! INTERWEBS JUST GO BOOM! OK, break out the walkie-talkies! Things I worry about more are mundane, prosaic stuff like the national debt going boom. Or inflation going boom. Or Taiwan going boom. Or flash flood from climate change making my house go boom. Nobody cares about these, the real-life, present stuff, because they aren't flashy things with Arnold as top headliner. Reply
Key considerations
- Investor positioning can change fast
- Volatility remains possible near catalysts
- Macro rates and liquidity can dominate flows
Reference reading
- https://www.tomshardware.com/tech-industry/artificial-intelligence/SPONSORED_LINK_URL
- https://www.tomshardware.com/tech-industry/artificial-intelligence/openai-and-anthropic-are-reportedly-investigating-tens-of-thousands-of-ai-security-incidents-openai-pauses-testing-after-ai-kill-switch-fails-to-stop-a-rogue-agent-report-says-problem-is-orders-of-magnitude-more-complex-than-what-is-publicly-known#main
- https://www.tomshardware.com/membership
- Virginia Tech lab 3D prints a liquid metal composite to guide heat, boost thermal conductivity 40x
- Walmart price drop slashes $461 off Gigabyte's RTX 5080-powered gaming laptop with 32GB of memory
- Novel attack slashes computing power needed to crack textbook RSA cryptography
- How Open Science Can Help Researchers Prepare for the Next Pandemic
- ChatGPT-6 Astra cracks 85-year-old 1941 Enigma-coded message in two days
Informational only. No financial advice. Do your own research.