Frontier AI faces pricing reckoning as token volume explodes 25-fold — mid-tier models deliver 90% of flagship capability at one-sixth the cost

Frontier AI faces pricing reckoning as token volume explodes 25-fold — mid-tier models deliver 90% of flagship capability at one-sixth the cost

Ditching the cloud for local AI — how I use two mini PCs to process millions of tokens a day and save money on costly API fees

Kimi K3 rocks the AI industry as Moonshot AI undercuts closed-source American competitors on price

And even then, the top AI models are absurdly expensive compared to the models on the Pareto frontier.

Claude Fable 5.1 comes with a 75% cut in the cost of its cache write pricing, and Artificial Analysis still clocked it at $3.69 per task on its Intelligence Index test. That comes from much more expensive answers and reasoning, because while Fable 5.0 has more expensive cache write costs, it's $3.14 per benchmark task. But that's 50% more expensive than Claude Opus 5 on the same task, which is double again the cost of GPT 5.6 Sol.

Then costs really start to crater, especially when you consider the intelligence of the more affordable models.

Google's Gemini 3.8 Flash (high) is a powerful model, able to score a 59 on the Intelligence Index test. But it costs a mere $0.58 per task on the Index test – less than 1/6th the price of Claude Fable 5.1, with just a 10% drop in intelligence scoring. OpenAI's GPT 5.6 Sol (high) costs $0.43, with an intelligence score of 57.

Chinese competition is right there in the mix, too. The daunting Kimi K3 (max) can manage a 60 on the intelligence benchmark, with a per-task cost of $0.84, while its Kimi K3 (low) variant offers a 48 score on intelligence at just $0.24 per task. Deepseek V4 Pro is arguably one of the most impressive, with a 53 and $0.27, respectively.

At the time of writing, Meta's Muse Spark 1.3 (xhigh) holds the Pareto frontier title, with a score of 61 and a per-task cost of just $0.55. It stole that top spot from Google's Gemini 3.8 Flash, which wore the crown for just 3.5 hours.

Gemini 3.8 held a spot at the pareto frontier for *checks notes* 3.5 hours https://t.co/P1A46LAy1M September 2, 2026

The perspective and approach of the business community to AI use has been equally terrifying and fascinating. While we've all felt the fear of AI invalidating skills we've spent years acquiring, business leaders have swung massively between demanding AI use at a grand scale and then quickly following it up with, "oh god, no, not that much."

Uber famously blew through its annual AI budget in just a few months, and tokenmaxxing leaderboards saw one unnamed company eat through half a billion dollars worth of tokens in just a few weeks. But while everyone is certainly taking costs a lot more seriously than they once were, that's not slowing AI usage. Indeed, as more effective intelligence has become more affordable, token usage is exploding.

One of OpenRouter's engineers published a chart showing that overall paid token use had increased 25 times in the past year, and doubled over the past month alone.

very normal month of token growth nothing to see here pic.twitter.com/V2huOmNcYK August 31, 2026

This increase appears to be coming from some of those middle-of-the-pack, affordable intelligence models. According to OpenRouter's LLM rankings , the most used model for the past month was OpenAI's GPT 5.6 Luna, with close to 12 trillion tokens. With its intelligence score of 52 and a per-task cost of just $0.05, it's right on the Pareto line at the cheapest end of the spectrum.

Right behind it, though, is Chinese developer Z-Ai with its GLM 5.3 Flash. It's at 11.4 trillion tokens in the past month, a more than 1,000% increase month to month. Its intelligence-to-price ratio is 57 to $0.09. Deepseek v4 Flash is right there with it, and other Chinese, intelligent-enough but very-affordable models round out the pack.

In comparison, the major, expensive models are barely being used at all. Fable 5's monthly use is in the low billions of output tokens, and even OpenAI, with its massive user base, is only cracking 1.8T monthly tokens with its 5.6 Sol.

Besides the bonkers business model for many of those involved, there are intriguing patterns emerging in AI usage. People can find ways to use lots of tokens, but they are only willing to pay so much for them . They want intelligence at as low a price as possible, and there is a crossover point where one becomes more important than the other.

While cynics argue that benchmarks are gamed, and boosters are still heralding the coming of their AI savior, the actual economics of the industry paint a much clearer picture. Intelligence has a price, but it's much lower than some of the frontier model developers are able to build it for. As models become ever more efficient and the hardware for inference grows ever more powerful, we may reach a point where what large language models can do effectively is affordable enough that anyone can use it as much as they want.

What that means for the major companies who spent hundreds of billions of dollars to get us to that point, very much remains to be seen.

Jon Martindale Freelance Writer Jon Martindale is a contributing writer for Tom's Hardware. For the past 20 years, he's been writing about PC components, emerging technologies, and the latest software advances. His deep and broad journalistic experience gives him unique insights into the most exciting technology trends of today and tomorrow.

Key considerations

  • Investor positioning can change fast
  • Volatility remains possible near catalysts
  • Macro rates and liquidity can dominate flows

Reference reading

More on this site

Informational only. No financial advice. Do your own research.

Leave a Comment