Tokenomics: Why Making AI Pay Is Tricky

If you have ever used a free version of ChatGPT or one of its rivals, you were probably enjoying a high‑value service at no cost. Microsoft, Google, Anthropic and others have poured hundreds of billions into developing the large language models (LLMs) that power these offerings. While the premium versions add extra features for coding, business analytics or specialised assistance, the core model remains free for basic use.
For the companies behind the free layers, the challenge is turning investment into revenue. Paid plans are offered, but setting a price in an environment where token consumption fluctuates is far from straightforward. "Trying to tie someone into a cost model for the next 12, two or three years doesn’t make any sense, honestly, because we don’t know," explained Simon Gooch, a senior executive at Saviynt.
AI tokens are the building blocks of every LLM interaction. A prompt is first split into tokens, then the model responds with its own tokenised output. The length and complexity of both the prompt and response determine how many tokens are consumed.
Unfortunately, the relationship between token usage and model behaviour is not linear. Small changes to a prompt can yield vastly different answers, and different models may produce alternative replies. In agentic systems where multiple AI agents collaborate, token use rises further, compounding the uncertainty.
Goldman Sachs estimates token consumption could jump 24 times between 2026 and 2030 to reach 120 quadrillion per month, driven by a shift from isolated models to autonomous agents. Yet most companies and individuals cannot track how many tokens they burn until the bill arrives.
Microsoft is reportedly tightening the use of its engineers’ access to third‑party coding tools, while Uber appears to deplete its AI coding token budget in a single year. Even large providers rarely reveal how token costs translate into customer invoices.
Will Venters, a professor at the London School of Economics, highlights that organisations often lose sight of token economy until they run out of credit. “People are finding it really hard to manage that cost… it’s a non‑deterministic output, so it’s a non‑deterministic value,” he said.
Smaller firms are hiding behind flat‑fee personal accounts, which they admit could be problematic as giant vendors clamp down on usage. Oliver King‑Smith of smartR AI cautions that big AI platforms may force stricter controls once shareholder pressure mounts.
Venters stresses that businesses should choose models carefully and write precise prompts. “Don’t send someone in your family out to get the weekly shop without any detailed instructions as to what you expect in the basket, right?” he used to illustrate the need for clarity.
Implementing AI in products amplifies token costs further, as developers use tokens for testing, security checks, and guard rails. The benefits may outweigh the cost, but the non‑linearity makes budgeting challenging.
Bill Peterson of Sumo Logic stresses that the sector still debates how best to price AI services. He said the firm is "still having some fun conversations about this internally" and considers options such as price‑by‑result or bundled incidents, but warns that changes in LLM providers’ pricing could shift the landscape.
















