
When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works .
With all the advancements in models over the last two years, having a bot that answers questions of simple-to-moderate difficulty is now old news, as they all do that with reasonable accuracy. Agentic workloads are where the real potential is at, letting bots loose on multi-stage, repeatable tasks that would previously take a human days, if not weeks, to perform. Grant a bot access to your billing system, some Excel spreadsheets, and your CRM, and ask it something like "who are my profitable customers by category and what are their trends" becomes child's play.
While it's trivial for you to ask that question and get a fairly accurate answer back in a couple minutes, behind the scenes there is a lot of processing going on, and far more than you'd expect. AI computing time is measured in tokens — a short question and answer might take somewhere between 200 to 2,000 tokens, and one that requires the models to do some internet research might be around 1,000 to 4,000. An agentic task, though, can easily spend millions of tokens on a seemingly innocuous request. How? Token amplification, a recently coined term.
In a simplified manner, because a model has no memory or cognition, every time you ask it another question in a conversation, it will re-load and process the entire exchange — everything you wrote, everything the bot replied with, and every file you uploaded. That means that additional questions in a long conversation progressively get more costly. Each question in a conversation might only need 500 tokens by itself, but reprocessing all the previous information adds up, so the second one might spend 600, and so on and so forth. The conversation as a whole uses up the cumulative number of all individual interactions, and as an added penalty: Response speed also tends to get slower as chats drag on.
The aforementioned task of generating a report will have to be run in stages, say three for looking up Excel sheets, four for the CRM, perhaps a half-dozen web searches for contextual information about products, and a good dozen intermediary processing and calculation steps. Each step tacks on potentially several thousands of tokens, and by the end of it, you may be looking at millions of tokens cumulatively spent for all the steps combined.
AI costs begin to bite as agents may increase token demand by 24 times, says Goldman Sachs report
Talent over tokens: AI models are becoming more expensive to run, and productivity gains are limited
OpenAI CEO Sam Altman admits AI token costs are becoming 'a huge issue.'
Key considerations
- Investor positioning can change fast
- Volatility remains possible near catalysts
- Macro rates and liquidity can dominate flows
Reference reading
- https://www.tomshardware.com/tech-industry/artificial-intelligence/SPONSORED_LINK_URL
- https://www.tomshardware.com/tech-industry/artificial-intelligence/the-enduring-paradox-of-the-ai-economy-models-get-better-and-more-efficient-yet-costs-can-still-easily-spiral-out-of-control#main
- https://www.tomshardware.com/my-account
- TSMC confirms significant yield and performance improvements in A14 update — strong interest from AI/HPC and smartphone customers
- Joyride Through July With 12 Games Coming to GeForce NOW
- New plugin unlocks granular VRAM temperature tracking on Nvidia RTX 50-series GPUs — community cracks open Blackwell's forbidden telemetry sensors
- Shrink your PC setup with up to 25% off on Kamrui mini-PCs — big savings on the H1 for gamers and the Hyper H2 for pros
- Florida man arrested after allegedly stealing $220,000 in crypto using malware hidden in Steam Games — 8,000 devices infected
Informational only. No financial advice. Do your own research.