The Token Trap
A Tale of Two Takeout Orders
The Quick Version
Executives are treating token usage as a scoreboard, but a token is a unit of expenditure, not value, and high consumption is a cost to interrogate rather than a trophy to celebrate.
The same box of General Tso’s chicken runs $15 at the local specialist or $50 at the Cordon Bleu kitchen, and in AI that 3x gap buys conversation, infrastructure, and timing, not a better outcome.
The switch that fixes this flips the question from how much AI you are using to how much enterprise value you extract per dollar, and there is a scorecard and an architecture that make the move measurable this quarter.
I was having coffee in a hotel lobby when I overheard a conversation between two investors. They were playing a game of one-upmanship about whose portfolio company was consuming more tokens, talking about usage per developer. These were tokens with a B—as in billions.
That is a problem.
As companies rush to implement AI, a misleading vanity metric of the moment has taken hold: token usage. Executives look at dashboards showing millions of tokens consumed, treating volume as a proxy for innovation or digital transformation. Trade press coverage treats tokens like a new global currency, a digital oil fueling the modern corporation.
This misunderstanding has paved the way for the latest performative corporate trend: token-maxing. In a bid to climb internal AI-adoption leaderboards, employees are now inflating their usage, running endless agent loops and generating redundant drafts just to look productive on paper. The fallout is already public: reports of major enterprises like Uber burning through an entire annual AI budget in the first 4 months of ‘26, after runaway token consumption blew past forecast.
This confusion about what a token actually represents is driving bad decision-making. By focusing entirely on consumption, leadership is ignoring the only metric that matters: the actual business value generated by those tokens.
Some basics: AI models speak in tokens. Every token is roughly 3/4 of a word. Every word we input, the datasets we upload, and the internal reasoning steps the models generate are all tokens. How effective those tokens are, and what each one costs, depends on user prompts, the choice of large language models (LLMs), the data centers where those models sit, and a variety of other factors.
Referring to raw token volume as a measure of progress is like measuring the success of a legal department by the number of pages they print. It measures activity, not productivity. A token is not a unit of value. It is a unit of expenditure. High usage is not an achievement. It is a cost. To understand the true ROI of AI, we must stop measuring computational activity and start measuring enterprise outcomes achieved per dollar spent.
To see how this plays out, let’s step away from the data center and order some takeout. Personally, I’m in the mood for Chinese food.
A Tale of Two Takeout Orders
Imagine you have a craving for Chinese food. You have two options for takeout, and both will result in a virtually identical product: a box of General Tso’s chicken.
Situation One: The Local Specialist
You walk into your local, takeout-only Chinese restaurant. It is tucked away on a side street in a low-rent neighborhood. Inside, you see the same woks, knives, and bowls they have used for 10 years. They source their ingredients from a local supply house in Queens.
You walk up to the counter and say, “Number 17.” The cashier doesn’t say a word. He turns and yells a single word into the back. The cook yells back to confirm. Inside the kitchen, the operation is lean. The cook asks for “chicken,” an assistant grabs it, and the chef says, “Box it.” A takeout container, already laden with rice, is filled, bagged, and placed on the counter. The cashier tells you the price: $15.
Better yet, if you come in during the middle of the day, they have a lunch special: the exact same meal for $10. Because they know the demand is predictable and they can prep in bulk, they pass that operational efficiency on to you.
In AI terms, this is a Small, Fine-Tuned LLM. It is the digital equivalent of that low-rent side street restaurant. The infrastructure (the data center) is optimized for cost, the hardware is stable, and the ingredients (the training data) are specific and local. The lunch special represents batch processing or reserved capacity, situations where you can get the same high-quality output for a fraction of the cost by processing your data during off-peak hours.
Everyone involved spoke very little (low token usage), and what you were charged for those words was incredibly low (low cost per token).
Situation Two: The Cordon Bleu Kitchen
Now, imagine you go for takeout at a fancy, high-end Chinese restaurant with cloth napkins and an in-house sommelier. This place sits on a main thoroughfare, a massive avenue in a high-rent district. They source all their ingredients from boutique farms upstate, and they replace their kitchen equipment every year with the latest and greatest, including ultra-expensive, high-tech rice cookers.
The experience begins with a warm greeting. The maitre d’ asks how your day was and inquires about your preferences. You order the General Tso’s. “Wonderful, sir,” he says. “That will take about 20 minutes. Is that okay?”
He writes a detailed note and hands it to a waitress. She takes it to an assistant chef, who takes it to the Primary Chef. This chef is a master of thousands of cuisines. He trained at a Cordon Bleu restaurant. He looks at the order and begins to overthink it. He turns to the waitress: “Are you sure he wants General Tso’s and not General Sue’s chicken?”
The waitress isn’t sure. She goes back to the maitre d’, who asks you for clarification. You confirm your order. The message travels back through the chain. The Primary Chef then debates the best preparation method with his staff, calls his brother in San Francisco (who is also a chef) for his opinion, asks assistant chefs to prepare artisanal components, and finally produces the meal. It is placed in a heavy-duty box, which goes into a branded paper bag. Inside, they’ve added heavy-duty forks, linen-feel napkins, and a complimentary set of chopsticks you didn’t ask for.
The bill? $50.
This is a Massive Frontier LLM. It is brilliant, polymathic, and capable of almost anything. But because it sits on high-rent infrastructure (H100 clusters) and uses the latest and greatest equipment (cutting-edge R&D and training techniques), the cost of its existence is passed directly to you. It over-communicates, over-thinks, and over-packages because it was built to handle world-class complexity, even when you just want a quick snack.
The Outcome: The Cost of Conversation
When you get home and open both bags, you have the exact same chicken. In fact, you might find that the $10 lunch special tastes better because the local cook does nothing but make that specific dish all day.
The ultimate business outcome is identical: a satisfied customer eating lunch. But the path to that outcome represents a 5x cost differential.
The $40 difference in price wasn’t for the food. It was for the conversations, the infrastructure, and the timing. Each word spoken in that fancy kitchen represents a token. Each of those people, the maitre d’, the waitress, and the master chef, has a higher salary (unit cost) because they are operating in an incredibly expensive environment. When evaluating AI, the question is not “how much did we say?” but “did we get the chicken to the customer for $15 or $50?”
The Moving Target: Tokens are Not a Constant
Here is where the currency myth truly falls apart: the price of a token is not fixed. Who you use to provide them, how you use them, and when you use them, matters.
Imagine if the fancy restaurant changed its prices based on the time of day. If you walk in at 7:00 PM during the dinner rush, the maitre d’ might charge you double because the kitchen is at peak capacity. If you go on a Tuesday morning, it might be half price. Furthermore, if you go to a different fancy restaurant down the street, they might have a different master chef with a different per-word rate entirely.
In AI, a million tokens spent on a Monday might cost twice as much as a million tokens spent on a Tuesday due to provider price drops, time-of-day spot pricing for compute, or simply because you switched from one LLM provider to another.
If you treat tokens as a currency, your balance sheet will look like a rollercoaster. You spent the same amount (1 million tokens), but the cost varied by 50%. Without knowing the outcome, the quality of the chicken, you have no way to determine if that was a good expenditure or a total loss.
Measuring the Right Metrics
To manage this complexity, businesses must move away from vanity metrics. Instead of counting raw tokens, leadership must tie token expenditure directly to tangible business outcomes. If you cannot tie token consumption to a closed customer ticket, a processed invoice, or a generated lead, you are simply subsidizing expensive digital chatter.
Here are three specific KPIs that track actual financial and operational performance:
Total Token Expenditure (TTE): This measures aggregate spend rather than volume. It answers the question, “How much money did we actually spend at the restaurant?” and treats AI as a budget line item that must be justified by business value.
Blended Cost per Token (BCPT): This tracks vendor pricing shifts, time-of-day premiums, and provider competition. It allows you to see if your AI costs are rising because you are using more AI, or simply because you are ordering during peak hours or from more expensive restaurants.
Model-Specific Cost per Outcome (MCPO): Instead of just looking at utilization, calculate the cost of the AI infrastructure required to solve a specific problem. If a fine-tuned model solves a support ticket for $0.02, while a frontier model solves it for $0.50 with the exact same customer satisfaction rating, the fine-tuned model is the clear winner.
The Efficiency Leaks: Where the Money Goes
When you look at your AI bill, you are often paying for infrastructure leaks that provide zero value to the end user:
Prompt Bloat (The Unwanted Utensils): System prompts filled with redundant instructions add to your token count without adding to the flavor. You are paying for packaging you’re going to throw away.
The Harness Penalty (The General Sue Debate): Internal reasoning steps and loops generate thousands of tokens that the user never sees. This is the AI thinking out loud, and you are paying for every syllable of that internal debate.
Model Mismatch (The Cordon Bleu Chef): Using an expensive, high-intelligence model for a low-intelligence task is the fastest way to destroy your ROI.
The Strategic Solution: Least-Cost Routing
The future of business AI relies on model routing or least-cost architectures:
The Triage Agent: A very low-cost model evaluates the incoming request.
Complexity Scoring: Is this a standard “Number 17,” or is this a unique request for a complex fusion dish?
Dynamic Dispatch: Simple tasks go to pennies-per-million models or lunch special batch queues. Only high-value synthesis tasks go to the master chef.
Compression: High-rent models receive only the most vital information, so every expensive token carries high value.
Conclusion: From Usage to Outcomes
A report stating that “10 billion tokens were consumed this quarter” is merely reporting the volume of gas burned, not the mileage achieved, the freight moved, or the actual cost of the fuel. In a business context, a high token count should be viewed with suspicion until it is correlated with a specific ROI.
The goal of an AI strategy shouldn’t be to maximize token usage. It should be to maximize value while minimizing the conversation cost and rent required to get there. The conversation must change from “How much AI are we adopting?” to “How much enterprise value did we extract from our compute spend?” Let’s stop celebrating how much we eat, and start measuring how well we are fed.
In the end, the most successful AI winners will look a lot like that local takeout shop: quiet, efficient, specialized, and focused entirely on putting the food in the box for the lowest possible cost.
N.B. When I got home I checked my token usage for the month. A measly 320mm token. I better get back to work
Key Takeaways for Busy Leaders
Save/Revenue: Price your AI on cost per outcome, not token volume, and matching the model to the task can deliver the identical result for up to 5x less.
Pitfall: Token-maxing and runaway agent loops can torch an annual AI budget in months, as Uber’s early-2026 overrun showed, so treat a climbing token count as a cost to investigate, not a win.
Deeper Dive: The 3-KPI scorecard (TTE, BCPT, MCPO) and the least-cost routing blueprint above are your team’s token-economics reference, save this issue and run them against your current stack this quarter.



