Flexera logo
Image: Tokenmaxxing to valuemaxxing: What every IT and finance leader needs to know (2026)

Over the last two years, the enterprise playbook for AI adoption was simple: use more of it. More tokens, more seats and more agentic workflows, on the belief that heavier usage would lead to faster returns. In 2026, that approach picked up a name, tokenmaxxing, as some companies began using token consumption and token leaderboards to measure and encourage AI adoption. But as AI usage grows, so does the danger of mistaking consumption for productivity and business value. That danger is driving the shift from tokenmaxxing to valuemaxxing. 

Valuemaxxing measures AI by what it delivers, such as tasks completed, hours saved and costs avoided, rather than by how much it gets used. For IT leaders trying to keep AI spending under control and finance leaders trying to explain that spending to the board, understanding the difference is a must. 

In this article, we’ll break down what tokenmaxxing and valuemaxxing mean, why enterprises are shifting from usage-based metrics toward outcome-based measurement and how to apply that thinking to AI adoption, AI cost management and ROI. 

Is the AI cost crisis real? 

Yes, the AI cost crisis is real. So what is it? The AI cost crisis is the widening gap between what enterprises spend on AI and the measurable business value that spending produces. AI budgets are growing faster than almost any other IT line item, but visibility, governance and proof of return lag behind. 

Start with the spend. Gartner’s latest forecast, published September 16, 2026, puts worldwide AI spending at about $2.7 trillion in 2026, up 49.5% from 2025. That’s higher than the $2.59 trillion Gartner projected in May, and the growth rate is more than three times the 14.2% Gartner expects for IT spending overall. 

 The value side looks murkier and thinner. A July 2025 report from MIT’s Project NANDA found that 95% of organizations studied saw no measurable profit-and-loss impact from their AI initiatives. Newer data points the same way. In KPMG’s Global AI Pulse Q2 2026 survey of 2,145 senior leaders, only 7% said they’d reached the stage KPMG calls established ROI. 

Visibility isn’t much better. Our Flexera 2026 State of ITAM Report puts a finer point on it: just 31% of organizations say they have accurate visibility into AI software. And 59% say wasted AI software spend rose over the past year, the highest of any software category in the survey. 

Taken together, those numbers describe a familiar pattern for anyone who lived through the early years of cloud cost management: spend scales quickly, oversight lags behind and the bill lands on finance’s desk with little explanation. For a closer look at where AI spend goes, see our analysis of how AI is reshaping cloud and data cloud spend. And a good chunk of that bill traces back to one habit: tokenmaxxing. 

What is tokenmaxxing? 

Tokenmaxxing is the practice of pushing people to use as many AI tokens as possible and treating that consumption as proof of adoption or productivity. IBM defines tokenmaxxing as incentivizing employees to maximize token usage, usually to encourage experimentation with AI tools. 

To see why it caught on, start with the unit itself. An AI token is a piece of text that a large language model (LLM) processes as input or generates as output. A token can represent part of a word, a whole word, punctuation or other text units. For typical English text, one token works out to about four characters, or roughly three-quarters of a word. The exact token count varies by model, tokenizer and language.  

Tokens also provide a convenient way to measure AI usage. Many AI model APIs charge based on input and output tokens, often at a per-million-token rate, though pricing and billing units differ across models and providers. That makes token usage easy to measure and, by extension, tempting to use as an AI adoption metric. 

The “maxxing” suffix is internet slang for pushing one thing to an extreme. Applied to AI, tokenmaxxing means pushing token consumption as far as possible, especially when teams treat higher usage as proof of AI adoption or productivity. 

The tokenmaxxing trend peaked around spring 2026, as companies made AI usage more visible through internal dashboards, leaderboards and adoption metrics. In March, Nvidia CEO Jensen Huang said he’d be ”deeply alarmed” if a $500,000-a-year engineer didn’t consume at least $250,000 worth of tokens. His remarks became a defining example of the usage-first mindset behind tokenmaxxing.  

Around the same time, Uber ranked teams by AI tool usage and then burned through its entire 2026 AI coding tools budget in four months. Other tech giants ran internal token leaderboards of their own. 

Inside most enterprises, tokenmaxxing doesn’t always look that extreme. It can show up through practices such as: 

  • Giving employees broad access to AI tools without clear usage policies, cost controls or ownership 
  • Defaulting to the largest or most expensive AI model for tasks a smaller model could handle 
  • Running long-lived AI agents or parallel agent workflows without limits on retries, tool calls or execution time 
  • Tracking token consumption as an adoption metric instead of measuring completed work, quality or business outcomes 

The problem gets clearer when a usage metric turns into a target, a classic case of Goodhart’s law: when a measure becomes a target, it stops being a good measure. That’s what happened in 2026. Some employees at big tech companies inflated usage numbers by directing AI at pointless tasks just to meet targets, and the backlash came quickly.  Within months, several large companies retired token leaderboards, capped routine token use or issued guidance to curb token spend. 

To be fair, tokenmaxxing wasn’t all bad. It got skeptical teams to try the tools, and some of that experimentation paid off. The trouble started when the adoption metric quietly became the success metric. 

 If you already know where your token spend is going, our breakdown of prompt caching covers one practical way to reduce repeated input-token costs. Cutting spend is only half the job, though. Even honest token counts say very little about value. 

Why token counts make a poor scoreboard for AI value 

Token counts fail on two fronts, and that’s why tokenmaxxing breaks down: they don’t tell you how much value the work created, and they don’t reliably tell you what that work cost.  

Start with value. AI tokens measure activity, not outcomes. A high token count tells you a model processed a lot of text. It doesn’t tell you if the output was useful, accurate or worth the spend. 

AI cost isn’t much clearer, because token counts lump together very different charges: 

  • Output tokens usually cost several times more than input tokens, while cached input is often billed at a steep discount 
  • Reasoning models bill their hidden “thinking” as output tokens, even though it never appears in the answer 
  • Tokenizers vary between models, so switching models can change your token count even when the task and the text stay the same  

AI agents multiply all of it. Gartner estimates that agentic models need 5 to 30 times more tokens per task than a standard chatbot, and it predicts inference costs per agentic workflow will rise more than fivefold through 2028, even as unit prices fall. 

So a token count alone can’t tell you if AI spend is healthy. You need the other half of the equation: what the work produced. That’s where valuemaxxing comes in. 

What is valuemaxxing? 

Valuemaxxing is the practice of judging AI by the business outcomes it produces, such as tasks completed, hours saved, revenue gained or costs avoided, instead of by how much AI gets used. It’s the direct response to tokenmaxxing’s core flaw: consumption and value aren’t the same thing. 

Instead of counting tokens, valuemaxxing asks: 

  • How many tasks got done, and at what quality? 
  • How much time did we save, and where did that time go? 
  • How much rework, risk or spend did we avoid? 
  • What did each good outcome cost us? 

The term is generally credited to Marc Boroditsky, chief revenue officer at AI infrastructure company Nebius, and has since been widely adopted across enterprise technology commentary. Our own Chief Product Officer, Becky Trevino, wrote about valuemaxxing as a reset for how enterprises manage technology spend. Her post lays out three moves for leadership: 

  • Create structured oversight, with AI-focused teams and shared metrics that govern technology investment with the same rigor human resources applies to talent 
  • Eliminate application redundancy, starting with a complete inventory across on-premises, SaaS and cloud environments 
  • Ditch technology shortcuts like vibe coding, which can trade reliability, security and long-term maintainability for short-term speed 

Here’s tokenmaxxing vs valuemaxxing, side by side: 

Comparison of enterprise AI cost outcomes under tokenmaxxing vs valuemaxxing - What is Tokenmaxxing - What is Valuemaxxing - Tokenmaxxing to Valuemaxxing  

Comparison of enterprise AI cost outcomes under tokenmaxxing vs valuemaxxing

One thing valuemaxxing isn’t: a token diet. IBM points out that token minimization falls into the same trap as tokenmaxxing, because both treat the token count as the number that matters. Strip out too much context and the costs don’t vanish. They move into retries, extra reasoning, more tool calls and human rework. 

For a fuller breakdown of how to put valuemaxxing into practice, read our guide, Valuemaxxing: Why it matters and how leaders can master it. 

Why tokenmaxxing became the default, and why it no longer works 

If valuemaxxing makes so much sense, why did tokenmaxxing become the default? Two reasons. First, early pricing made experimentation cheap: seat-based licenses capped what most employees could cost, promotional credits often covered early pilots and per-token prices kept falling. Second, AI usage was easier to measure than outcomes. Token counts come straight off the invoice, while time saved or revenue gained takes real work to prove. 

 It stopped being enough once usage started showing up as unplanned, unexplained AI costs. Our 2026 State of ITAM Report found that 59% of organizations reported an increase in wasted AI software spend year over year, and that just 29% currently measure the value of the AI software they’ve bought at all. That’s the gap that lets tokenmaxxing run unchecked. 

 Cost pressure is already forcing changes. In KPMG’s Q2 2026 Global AI Pulse, 49% of organizations had scaled back, delayed or paused AI agent deployments when expected costs began to outweigh anticipated value, a trend we explore in what enterprises are learning as AI budgets balloon. Unsanctioned shadow AI tools add another layer of AI spending that finance often can’t see until renewal. 

 Read more: How to detect and govern shadow AI. 

 Analysts expect AI governance to catch up. IDC predicts that by 2028, 70% of leading AI-driven enterprises will dynamically route tasks across multiple models based on cost and performance, rather than defaulting to the most powerful (and most expensive) option for everything, an approach Flexera explores in more technical depth in its guide to agentic FinOps for AI. 

How to move from tokenmaxxing to valuemaxxing 

Governance sets the rules. These six valuemaxxing techniques improve your unit economics by lowering the cost of each outcome without taking the outcome away: 

 

Technique  What should you do about it  And why it helps 
Model routing  Send each task to the smallest model that meets the quality bar  Smaller models cost a fraction of frontier models per token 
Prompt caching  Reuse repeated set of instructions, tool definitions and documents across calls  Major providers bill cached input at a massive discount 
Batch processing  Run work that can wait, such as nightly classification or evaluations, asynchronously  Batch jobs commonly cost about half the standard rate 
Context and output limits  Trim stale context, cap output length and load only the tools a task needs  Agents resend context at every step, so waste compounds 
Budgets and kill switches  Give every team, app and agent a token or dollar budget, with anomaly alerts  KPMG found only 40% of organizations have usage or token budgets in place 
Tool rationalization  Consolidate duplicate AI subscriptions and retire idle licenses  Overlapping tools and idle seats add cost without adding value 

 

Just don’t cut so hard that the model loses the context it needs. That’s the token minimization trap, and the cost comes back as retries and rework. Some tasks don’t need an LLM at all, either. A rule-based workflow can handle deterministic jobs without spending a single token. 

Looking for additional techniques? Our breakdown of why token bills are becoming the new cloud bill goes deeper. But techniques only work when someone owns them, and that’s where many organizations get stuck. 

Who owns AI cost decisions: IT, finance or both? 

Both, but not for the same decisions. That’s the part most AI governance guidance skips. 

Our own 2026 State of ITAM Report found that AI governance responsibility is now split across IT, security and finance teams, with no single owner, and it stopped at recommending “shared responsibility” without saying who’s actually responsible for what. That ambiguity is itself a cost driver: when nobody owns a decision, it defaults to whoever moves fastest, usually the team closest to tokenmaxxing. 

KPMG’s data shows why clarity pays. In its Q2 2026 Global AI Pulse, only 24% of organizations named the CEO or executive committee as ultimately accountable for AI-informed decisions. Organizations with clearly defined accountability reported established ROI at more than 3x the rate of those without it (14% vs 4%). That’s a correlation, not proof of cause, but it’s a hard one to ignore. 

Good AI cost governance is the difference between tokenmaxxing by default and valuemaxxing by design. Here’s what a workable starting split looks like: 

  • IT owns model access, routing rules, usage monitoring and, with security, shadow AI detection 
  • Finance owns budgets, forecasts, cost allocation and, with procurement, contracts and renewals 
  • Both share decisions about what counts as value for each use case and when to scale, pause or retire a pilot 
IT and finance ownership split for AI cost governance decisions - What is Tokenmaxxing - What is Valuemaxxing - Tokenmaxxing to Valuemaxxing 

IT and finance ownership is split for AI cost governance decisions

 

This isn’t an official standard, so adapt it to your structure. It gives IT and finance leaders a starting point for a conversation many enterprises are having right now without one. One rule makes it stick: give every decision a single accountable name, even the shared ones. If you already run FinOps practices for cloud spend, let that team referee. 

How do you measure AI value? 

Once owners are in place, you need a way to measure AI value. Unit economics, such as cost per prompt, per image or per interaction, tell you how efficiently you run AI. Our guide to FinOps practices for AI covers that side of the equation. But unit costs can’t tell you if an initiative is worth running at all. That’s a separate question, and it’s the one valuemaxxing tries to answer. 

Four-category scorecard for measuring AI value: time saved, quality, revenue impact and risk - What is Tokenmaxxing - What is Valuemaxxing - Tokenmaxxing to Valuemaxxing  

Four-category scorecard for measuring AI value

 

A simple way to start is to score each AI initiative against four outcome categories, rather than on its token or compute cost alone: 

  • Time saved: hours of manual work removed or sped up 
  • Quality or error reduction: fewer mistakes, less rework, faster resolution 
  • Revenue or cost impact: direct financial effect, positive or negative 
  • Risk or compliance value: exposure reduced, audit readiness improved 

 

As an illustration, not a real case study, imagine an internal support chatbot: high on time saved (deflecting routine tickets), medium on quality (mixed accuracy on complex queries), low direct revenue impact, and medium risk value (reduces reliance on manual, error-prone processes). Scored this way, a relatively cheap tool might turn out to be low value, while a more expensive one earns its cost several times over. Neither conclusion is visible from the spend line alone. 

This is deliberately lightweight. The point isn’t to build a perfect model on the first attempt. It’s to force the value conversation before renewal, not after the AI budget has already ballooned. It’s also a practical way to estimate AI ROI one initiative at a time. When you’re ready to take the numbers to the C-suite, our guide to building a business case for AI cost management shows how to frame them. 

Where to start 

Here’s how to start valuemaxxing this quarter: 

  1. Build one shared view of AI spend between IT and finance, so both sides work from the same numbers before the value conversation starts
  2. Assign ownership using the split above, even as a first draft, rather than leaving AI cost governance to whoever’s closest to the spend
  3. Swap token targets for outcome metrics by retiring token leaderboards and agreeing on one outcome metric per use case
  4. Score initiatives against outcomes, not just AI usage, using the four categories above or a version adapted to your organization 
Quick decision guide for who owns an AI cost decision: IT, finance, or shared - What is Tokenmaxxing - What is Valuemaxxing - Tokenmaxxing to Valuemaxxing 

Quick decision guide for who owns an AI cost decision

 

Standards are catching up, too. The FinOps Foundation’s FinOps Open Cost and Usage Specification (FOCUS), often called FinOps FOCUS, normalizes billing data across providers, and version 1.5 is scoped to add AI model identity and input and output token consumption, with release targeted for December 2026. That will make a shared view of token spend across vendors much easier to build. 

For a more detailed, step-by-step implementation approach once ownership and metrics are in place, see our practical guide to FinOps for AI and our eight steps to managing AI costs and resources. 

Conclusion: From counting tokens to counting outcomes 

Shifting from tokenmaxxing to valuemaxxing is not a single, one-time initiative. Rather, it is a change in how you measure and govern your AI investments. For most IT and finance leaders, visibility is the starting point: know where AI spend goes, who owns it and what it produces. Once you have that, the tougher discussions about model routing, prompt caching, budgets and retiring pilots become far more manageable. 

Tokenmaxxing made AI easy to count. Valuemaxxing makes it worth counting. 

Explore how AI Cost Management in Flexera One gives IT and finance one shared view of AI spend. Or read the Flexera 2026 State of ITAM Report for the full data behind our AI visibility and waste findings. 

 

What’s the difference between tokenmaxxing and valuemaxxing?

Tokenmaxxing measures success by how much AI is used: tokens consumed, API calls made, and compute burned. Valuemaxxing measures success by what that usage actually achieved: tasks completed, time saved, costs avoided. The shift matters because token volume was never a reliable proxy for business value, and enterprises that keep measuring usage instead of outcomes tend to see AI costs rise without a corresponding rise in results.

Is tokenmaxxing always a bad strategy?

Not inherently. Heavy AI usage during early experimentation can be a reasonable trade-off while teams learn what works. It becomes a problem when that consumption-first approach persists after adoption matures, with no metrics in place to show whether the usage is delivering value.

Who coined the term valuemaxxing?

The term is generally credited to Marc Boroditsky, chief revenue officer at AI infrastructure company Nebius, and gained wider traction through coverage from outlets including IBM and TechTarget in 2026.

Does valuemaxxing mean using less AI?

Not necessarily. Valuemaxxing is about aligning usage with outcomes, not capping it. In some cases, that means using AI more, in a more targeted way; in others, it means cutting back where usage isn’t tied to a measurable result.

Who should own AI cost decisions, IT or finance?

Neither owns all of it. A workable split has IT owning technical guardrails and usage monitoring, finance owning budget thresholds and cost allocation, and both sides sharing decisions like what counts as “value” for a given use case and when to scale or retire a pilot. Flexera’s own research found that most organizations haven’t formalized this split yet, which is itself part of the problem. 

How do you measure the value side of valuemaxxing, not just the cost side?

By scoring initiatives against outcome categories, such as time saved, quality or error reduction, revenue or cost impact, and risk or compliance value, rather than relying on usage or unit cost alone. This is a different exercise from unit economics (cost per prompt or interaction), which measures efficiency but not whether the initiative is worth running in the first place. 

What role does model routing play in valuemaxxing?

Model routing means matching each task to the most cost-appropriate AI model rather than defaulting to the most powerful one every time. IDC forecasts that by 2028, 70% of leading AI-driven enterprises will manage this kind of routing dynamically, treating it as a core part of AI cost governance rather than an afterthought.

What tools help enterprises move from tokenmaxxing to valuemaxxing?

Platforms built for AI cost management, such as Flexera’s AI Cost Management solution, are designed to give IT and finance a shared, real-time view of AI spend across apps, agents, models, and infrastructure, which is the visibility most organizations currently lack.

 

          Flexera 2026 State of ITAM Report: How leaders are balancing AI cost optimization and governance

Sources: