
When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works .
AI is currently used by AI 5x more than it is used by humans. That number will accelerate to 10x and then higher and higher.We keep speaking to human adoption when trying to determine ROI, but the utilization and scale is exponentially larger than that. September 30, 2026
OpenRouter is “a leading AI model gateway and routing platform.” A chart from its head of insights, Peter Walker, lists “7-day average token usage on OpenRouter split by type.” It states that, since the February crossover point, agents are using 14x more tokens while human usage is up 2.8x. The a16z chart in Newman’s post shows the same data. OpenRouter sorts each API key into one of three categories: agentic, mixed, or human, using a “7-signal weighted composite score that includes inputs such as tool call rate, turn count, gap timing, and others.”
The mixed category, possibly covering behavior that is part agent and part human, grew 4.7x over the same period, according to our math. Depending on how that traffic splits, the agents’ lead over humans may vary. The data also measures token volume, not spending. This is data from only one platform, and the trend isn’t completely consistent, with dips in April and July. However, agent use is growing elsewhere. In McKinsey’s 2026 State of AI survey, 40% of respondents from large organizations reported scaling AI agents, up from 27% a year earlier.
Cached tokens also account for nearly all of the relative growth in token usage, a16z wrote. They cost far less than processing a prompt from scratch, but they still have to be held in memory, and a16z, an OpenRouter investor, ties that to rising demand for high-bandwidth memory (HBM). Models keep that stored context in what’s called the KV cache, and “the KV cache is outgrowing GPU HBM capacity,” according to our reporting .
The same pattern shows up in the logs of a call center consultancy that tested DeepSeek on rented Nvidia H200s. In its own agents’ September usage on Claude Code, “96% of all input was re-reading old conversation.” Coupled with OpenRouter’s data, this suggests the token count overstates the bill, but the hardware cost for memory remains very real.
Token amplification and the enduring paradox of the AI economy
Key considerations
- Investor positioning can change fast
- Volatility remains possible near catalysts
- Macro rates and liquidity can dominate flows
Reference reading
- https://www.tomshardware.com/tech-industry/artificial-intelligence/SPONSORED_LINK_URL
- https://www.tomshardware.com/tech-industry/artificial-intelligence/futurum-ceo-says-agents-use-ai-5x-more-than-humans-number-will-eventually-hit-10x-but-agents-are-mostly-rereading-what-theyve-already-seen#main
- https://www.tomshardware.com/membership
- Geekbench 7 results suggest OpenAI's dots run on nine-core AMD EPYC VMs
- Russian missile strikes take out popular piracy websites
- PewDiePie unveils ‘uncensored’ Ajax AI model for home PCs
- Why Deploying Physical AI at Scale Demands Safety at Every Layer
- How NVIDIA GPUs Help Accelerate OpenAI’s GPT-6 Astra Ultrafast
Informational only. No financial advice. Do your own research.