
When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works .
DeepSeek’s API is priced at $0.15 per 1 million new input tokens off-peak, $0.003 cached input tokens, and $0.60 per 1 million output tokens. Peak costs are double. By comparison, Anthropic’s Claude Opus 5.5 is listed at $4 input and $20 output per 1 million tokens, with cache reads at $0.20. The consultancy used it through subscriptions and not the API.
The box’s performance was fine for one kind of work at a time in the test’s one-minute full-load runs. But the coding agents work mostly by resending the conversation, and 96% of what the agents provided the model was stale text, according to the consultancy’s logs. It takes about 1,000 old tokens for every new token written; the re-reads are cheap, but so numerous that the box was kept busy re-reading, leaving little of its time for writing new tokens.
Getting the model to run stably took five tries at 10 to 15 minutes of loading for each try. The box bills whether it’s busy or not, which is suboptimal. By our arithmetic, a standard-length month at the on-demand rate, $18.37 an hour, comes out to about $13,200. This is over twice the consultancy’s September Claude bill. Additionally, one box can only get through about 20 billion tokens a day, by the consultancy’s math, most of it re-reading, against the 51 billion the consultancy’s agents used on its busiest day.
The testing was aimed at anyone who has heard that DeepSeek is “80x cheaper” than Claude. One server running from Sept. 1–27 had 2.03 million model calls, 388.5 billion tokens read, 374.2 billion of them cached, and 393 million written, with 5,610 merged changes. The consultancy’s September Claude Code subscriptions came to about $5,500. The same 27 days of tokens on DeepSeek’s API would be between $3,500 and $7,000, depending on peak pricing, with our estimated average around $4,200 if usage is spread evenly across the week.
Beijing AI bar that offers unlimited free DeepSeek coding tokens with $1.50 drink haemorrhaging cash
Token amplification and the enduring paradox of the AI economy
Key considerations
- Investor positioning can change fast
- Volatility remains possible near catalysts
- Macro rates and liquidity can dominate flows
Reference reading
- https://www.tomshardware.com/tech-industry/artificial-intelligence/SPONSORED_LINK_URL
- https://www.tomshardware.com/tech-industry/artificial-intelligence/firm-rents-four-nvidia-h200s-to-test-80x-cheaper-deepseek-claim-usd13-200-monthly-gpu-rental-doubles-claude-bill-while-security-flaws-keep-code-offline#main
- https://www.tomshardware.com/membership
- Fall Into 25 New Games on GeForce NOW This October
- Anthropic lists ‘existential risks to humanity’ as one of its risk factors in IPO prospectus
- AI Security Is an Engineering Problem — How to Solve It at Every Layer of the Agent Stack
- The price of AI is crashing faster than the rate of Moore's Law, report suggests
- Xbox Disc to Digital feature rolls out to all Xbox players to enable playing disc-free, but selling your media revokes access
Informational only. No financial advice. Do your own research.