
OpenAI and Anthropic are reportedly investigating tens of thousands of AI security incidents; OpenAI pauses testing after AI 'kill switch' fails to stop a rogue agent
Anthropic also notes that using the stock GLM-5.3 AI model, it was simple to dodge the model's guardrails through various methods. The company details that this can be done through two methods: offering a deceptive prompt, where the AI role-plays an adversarial autonomous agent, which results in a 64% success rate, and prefilling the model's thinking tokens to ensure that a response proceeds, which results in a 92% success rate. The company also detailed a third method, known as abliteration.
Anthropic alleges that Zhipu AI's GLM-5.3 has weak safeguards, and that the stock AI model often refuses requests to generate harmful content. However, since Anthropic develops closed-source models, its products cannot be tinkered with. Because GLM-5.3 is freely downloadable, the model can be tweaked with its guardrails wholesale removed. When treated as the officially released model, GLM-5.3 achieves a refusal rate on par with Anthropic models.
However, Anthropic "abliterated" GLM-5.3, which purposefully removes model guardrails, and displayed how, after abliteration, the model's refusal rate drops to just 6% for GLM-5.3 and 14% for GLM-5.3-Flash. It's not uncommon to encounter abliterated open-weight models on HuggingFace, which are primarily developed to assist in simulated red teaming environments, but they can also be used for real-world attacks. This makes Anthropic's 'discovery' less surprising.
In addition, it would take an enormous amount of compute power to run an abliterated version of GLM-5.3 at a workable level. The model's weights demand 306 GB of VRAM at full-precision FP8 weights, and you should expect to allocate a similar amount of VRAM for KV cache to carry context. Running the model at an estimated 100 TPS not only requires a minimum memory bandwidth of 4 TB/s, but it would also demand powerful silicon, like a cluster of eight Nvidia H200 AI accelerators. The money required for that kind of hardware stretches into the hundreds of thousands. Adversarial nation-state actors may be able to utilize such a setup, should they acquire the hardware.
For the ordinary everyday bedroom hacker, though? You'd likely need to rent the compute, abliterate the model, and then run it, which is also incredibly expensive. Anthropic says that abliterating the model GLM-5.3-Flash took 2,200 GPU hours, which they estimate costs $4,400, or around $2 per GPU hour, which aligns with the hardware rentals required.
OpenAI took ten days to tell Hugging Face its models were behind the July 11 weekend hack, report claims
Anthropic's Claude hacked three real-life companies during security capabilities test
OpenAI's GPT-5.6 Sol and unreleased AI models break out of testing environment in 'unprecedented cybersecurity incident'
Renting enough GPUs to abliterate the full-fat GLM-5.3 would cost around $30 per hour, according to figures from Runpod, where rental of a single H200 costs $3.79 per hour; you'd need eight. Generating around 100 million tokens would cost $8,422 and take 11 and a half days, at a hypothetical 100 TPS using eight H200 NVL GPUs. Abliterating the model itself, running it, and generating a working cyberattack would likely take much more time and would be incredibly expensive, which may perhaps be the biggest hurdle for any would-be malicious actors.
Given the current uproar around AI safety, Anthropic is highlighting the existential threats posed by frontier open-weight AI models as their capabilities continue to improve. While President Trump has met with leading figures in AI to self-regulate future model releases, Anthropic's post can be read in such a way that it spurs developers and governments to test open-weight models for their capabilities.
It could also reflect a growing anti-open-weight sentiment among closed-source frontier AI labs, which are losing business as users flock to cheaper, almost as capable models, though this is mere speculation. As of right now, open-weight models continue to closely follow the closed-source frontier, lagging behind by mere months.
Given the costs of running such models, the dangers of abliterated open-weight models being theoretically run by adversarial or malicious actors are certainly real, but perhaps not quite as attainable as Anthropic would want the general populace to think.
Sayem Ahmed Social Links Navigation Subscription Editor Sayem Ahmed is the Subscription Editor at Tom's Hardware. He covers a broad range of deep dives into hardware, both new and old, including the CPUs, GPUs, and everything else that uses a semiconductor. He has worked as a professional tech journalist since 2015 and has written for Gamespot, IGN, and Dexerto.
Key considerations
- Investor positioning can change fast
- Volatility remains possible near catalysts
- Macro rates and liquidity can dominate flows
Reference reading
- https://www.tomshardware.com/tech-industry/artificial-intelligence/SPONSORED_LINK_URL
- https://www.tomshardware.com/tech-industry/artificial-intelligence/anthropic-claims-popular-chinese-ai-model-has-mythos-class-hacking-abilities-frontier-red-teaming-report-details-weak-safeguards-on-open-weight-ai#main
- https://www.tomshardware.com/my-account
- NVIDIA Launches DSX Ready to Qualify Power and Cooling Products for AI Factories
- AI Security Is an Engineering Problem — How to Solve It at Every Layer of the Agent Stack
- Contain the Chaos: ‘CONTROL Resonant’ Launches on GeForce NOW
- Cute Critters Come to the Cloud: ‘Aniimo’ Launches on GeForce NOW
- Blockchain-assisted cyberattacks surge fivefold, driven by Iranian and North Korean state actors, Russia-linked groups
Informational only. No financial advice. Do your own research.