The Tokenmaxxing Problem: Why Token-Based AI Pricing Is Broken

On July 1, Palantir CEO Alex Karp told CNBC that “something has gone completely wrong” with how the industry sells AI. He was talking about token-based pricing, the model behind most frontier AI products, where the bill scales with how much the system is used rather than what that use produced. Two weeks earlier, Palantir had published a nine-point manifesto arguing that this model pushes enterprises to spend on volume instead of outcomes. Two weeks later, on July 14, Chamath Palihapitiya took the same argument to the same network. He put a number on it: at 8090, the company he named for a promise of 80 percent of the software at 90 percent less cost, token costs were doubling every 45 days while the resulting productivity gain sat close to flat. “My costs are doubling every 45 days; my upside is essentially flat,” he said.
Different companies, different vantage points, same complaint. Karp runs a company that has spent two years positioning against the frontier labs’ pricing model. Palihapitiya runs one that is currently paying it, on track to spend over $10 million a year on AI usage credits. Both concluded the same thing: enterprises have lost the ability to tell whether their AI spending is producing value, because the metric they are billed on, tokens consumed, was never a measure of value to begin with.
The word for this, tokenmaxxing, is new. The pattern behind it is not. Uber burned through its entire 2026 AI budget in four months, after ranking employees by usage on an internal leaderboard, and responded by capping spend at $1,500 per employee per AI tool. Meta ran a similar leaderboard, internally nicknamed Claudeonomics, that recorded 73.7 trillion tokens of employee consumption in a single 30-day window before CTO Andrew Bosworth wrote to staff that “token usage alone is not a measure of impact of any kind” and the leaderboard came down. Tesla capped employee AI spending at $200 a week in July, after engineers were burning through thousands of dollars weekly on their own. Three companies, three internal leaderboards built to reward usage, three reversals inside the same two quarters.
None of these companies pulled back on AI. All of them discovered that adoption and value are not the same number, and that a pricing model tied only to consumption has no way of telling them apart. A routine query and a genuinely hard one draw down the same token meter, because the systems processing them have no concept of which is which. Rewarding volume was always going to produce this outcome.
This is the exact problem matched-effort AI architecture exists to solve. An agent that sizes its own effort to the complexity of the task in front of it, minimal compute on routine work, deeper reasoning and a human on the cases that warrant it, does not have a tokenmaxxing problem, because it was never built to treat effort and outcome as unrelated in the first place. Call the alternative valuemaxxing: spend scales with what a case needs and what it is worth, not with how many tokens can be run through it.

That is the design principle behind Kapture’s AgentOS. The Execution Engine, Vitos and Command, is where the work happens. Vitos runs a ticket, chat, or call end to end: reading intent, pulling the account context, working the resolution, closing the loop with the customer. Command sits alongside it, pulling a person in exactly where a case needs judgment, a disputed charge, an angry escalation, a decision outside policy, and staying out of the way everywhere else. Every touchpoint gets handled. Not every touchpoint gets the same amount of agent effort, because a password reset and a fraud dispute were never the same task.
The Intelligence Engine, Calibrate and Pulse, is what keeps that judgment from going stale. Every conversation, ticket, and call is captured in real time as interaction data. SLAs, drop-offs, and escalations get recorded as operational signals. Calibrate audits that record continuously, not on a sample, and turns the patterns it finds into specific recommendations: what to do differently, where, and why. Pulse turns those recommendations into retrained agents and tuned workflows. Execution feeds intelligence. Intelligence sharpens execution. The loop runs on every interaction, not once at deployment.
That loop changes what token usage looks like over the life of a deployment, not just at a single point in time. Early on, cost per resolution runs higher, because agents are still calibrating against real cases rather than assumptions. As Calibrate flags what worked and Pulse feeds it back into the agents and the workflows around them, the same category of ticket needs fewer reasoning steps, fewer escalations, and less back and forth to close. Resolution quality goes up. Token cost per resolution goes down. That is the compounding advantage the loop is built for.
Pricing follows the same principle, even in the parts of AgentOS structured around usage. The target throughout is that cost per resolution falls as the system learns, not that it rises as volume grows.
Karp and Palihapitiya just put a name on an accounting problem enterprises have been living with for two years. The name changes nothing about what happens next. Enterprises that were already matching AI effort to task value are measuring the gap between spend and outcome today. The rest are about to start, the way Uber, Meta, and Tesla just did, after the bill arrived first and the explanation came after.
This is the architecture Kapture runs today inside BFSI and retail enterprises across India: agent effort matched to case value across lending, collections, and customer operations, not usage maximized across all of it.
Get in touch to see it in production.
Your Plan. Your Value. Your Growth.
Your business is different – and the pricing should reflect that.
Let’s build a plan that matches your goals, maximizes ROI, and scales with your success.






