
Based on the V4 generation , where Pro outscored Flash, better performance from the model family can be expected with DeepSeek V4.1-Pro, which currently does not have a firm release date. This may change the math in the future, but for now there are better ways to spend on tokens than renting GPUs to run an open model.
Follow Tom's Hardware on Google News , or add us as a preferred source , to get our latest news, analysis, & reviews in your feeds.
Shane Downing Social Links Navigation Contributing Writer Shane Downing is a Contributing Writer for Tom’s Hardware, covering consumer storage, PC hardware, and AI.
ndneubauer2 DeepSeek is “80x cheaper” than Claude. Well, obviously you have to compare API costs with API costs, not the cost of renting hardware. What even is the point of this article? Reply
Moosegoose This is the dumbest article I have read in weeks. This isn’t even AI slop, it’s devoid of intelligence. I hope the author doesn’t think this makes him some sort of authority on LLMs. “The box’s performance was fine for one kind of work at a time in the test’s one-minute full-load runs. But the coding agents work mostly by resending the conversation, and 96% of what the agents provided the model was stale text, according to the consultancy’s logs. It takes about 1,000 old tokens for every new token written; the re-reads are cheap, but so numerous that the box was kept busy re-reading, leaving little of its time for writing new tokens.” Do…. you not know how caching works? Or are you a fourth grader, because that is about the critical thinking proficiency demonstrated here. Also the company that rented 4xH200s 24/7 to serve themselves Deepseek to their coding agents…? Equally devoid of any brain cells. Probably taking advice from Shane Downing. Reply
loveday.ben Is this a paid advertisement for Anthropic? I'm sure if I rented a server with gpu from Azure Foundry and deployed a Claude model I'd have a similar experience… Reply
jsidhenneld Admin said: A call center consultancy rented four Nvidia H200s to run DeepSeek V4.1 Flash for Claude Code and found DeepSeek's API cheaper. Firm rents four Nvidia H200s to test '80x cheaper' DeepSeek claim : Read more This most stup1d articles i read from this website, not even ai slop just pure stupidity API cost vs API cost, NOT RENT VS API COST !!!!🤦 Reply
ryanga I am soooo relieved that these comments exist. I was shocked that the article didn't understand that getting sample drugs for free was not actually what drugs cost and why api costs are not the consumer costs if the companies are still running at a loss unless someone else builds and supports their infra Reply
machdisk So the firm were incompetent at setting up LLMs? 4xh200 and they only got 200 tok/s concurrent on dsv4.1 a 16B active in decode and 8B active in prefill model? How? Were they running llama.cpp in pipeline parallel or something? Thats a fraction of what they should have got Reply
Key considerations
- Investor positioning can change fast
- Volatility remains possible near catalysts
- Macro rates and liquidity can dominate flows
Reference reading
- https://www.tomshardware.com/tech-industry/artificial-intelligence/SPONSORED_LINK_URL
- https://www.tomshardware.com/tech-industry/artificial-intelligence/firm-rents-four-nvidia-h200s-to-test-80x-cheaper-deepseek-claim-usd13-200-monthly-gpu-rental-doubles-claude-bill-while-security-flaws-keep-code-offline#main
- https://www.tomshardware.com/membership
- PewDiePie unveils ‘uncensored’ Ajax AI model for home PCs
- Fall Into 25 New Games on GeForce NOW This October
- The price of AI is crashing faster than the rate of Moore's Law, report suggests
- Sony released the first CD audio player on this day in 1982
- How Open Science Can Help Researchers Prepare for the Next Pandemic
Informational only. No financial advice. Do your own research.