AI REVENUE AND GOVERNANCE
AI Unit Economics: The Margin Nobody Owns
93% of companies exceed their AI budgets and token cost varies 30x on the same task.

Your cost to deliver one unit of work now varies by up to 30 times. Your price does not vary at all. The difference between those two facts is your gross margin, and in most companies nobody is accountable for it.
Executive summary
- In a McKinsey survey of enterprise AI spending fielded in May 2026, 93% of respondents reported exceeding their AI budgets, and spend rises nearly fourfold as organizations move from isolated use cases to enterprise-wide adoption.
- Stanford’s Digital Economy Lab found that running the same task twice can differ by up to 30 times in total tokens, that human expert difficulty ratings only weakly predict actual token cost, and that frontier models systematically underestimate their own consumption.
- Taken together those findings say something specific: AI cost of delivery is not a number, it is a distribution, and it cannot be estimated by looking at the work.
- Most companies price flat over that distribution. A flat price over a 30x cost range is not a pricing model. It is a position.
- The structural problem is ownership. The CIO owns the spend line, the CRO owns the price line, and gross margin per unit of delivered work sits between them and belongs to neither.
- This is usually framed as a cost control problem. It is a pricing problem wearing a cost costume. Cutting AI spend by 20% while the price is wrong makes you 20% less wrong.
- And the cost cut itself usually underdelivers. Bain surveyed 951 companies and found the modal company targeted 11 to 20% savings and realized 0 to 10%. 44% said they plan to fund their next wave of AI investment out of savings from prior automation programs, which is to say out of money the same survey shows did not fully arrive.
The two findings that have to be read together
Each of these has been reported on its own. The combination is where the commercial problem lives.
Spending is out of control and largely invisible. McKinsey’s enterprise AI FinOps survey, fielded in May 2026 with 120 enterprise participants and 75 qualified respondents across five industries, found that 62% of organizations have moved past experimentation into active deployment, and that 93% exceeded their AI budgets. Spend rises nearly fourfold on the move from isolated use cases to enterprise-wide adoption. McKinsey’s own field experience puts 20 to 30 percent of AI spend as unaccounted for, fragmented across cloud providers, model vendors, embedded software features, experimentation environments and business unit purchases. Only 20 to 25 percent of companies have mature AI FinOps practices.
And the cost of any given piece of work is close to unpredictable. Stanford’s Digital Economy Lab analyzed eight frontier models on agentic coding tasks and found that runs on the same task can differ by up to 30 times in total tokens. Agentic tasks consume on the order of a thousand times more tokens than a simple chat exchange, because each action re-reads the accumulated context. Two further findings matter more than the headline:
- Task difficulty as rated by human experts aligns only weakly with actual token cost. You cannot look at a piece of work and estimate what it will cost to deliver.
- Frontier models fail to predict their own token usage and systematically underestimate it, with correlations reaching only about 0.39.
There is also a finding that should change how people think about premium models: higher token usage does not translate into higher accuracy. Accuracy tends to peak at intermediate cost and then flatten. Spending more does not buy a better answer past a point.
Why this is a margin problem and not an IT problem
Put those together from a commercial seat rather than a technology seat.
If you sell a product or service that AI delivers, your cost of goods sold now behaves like a variable input with a long tail, and you cannot forecast it from the shape of the work. Your average cost per transaction may be perfectly comfortable. Your ninety-fifth percentile may be underwater.
Most companies price flat. A subscription, a seat, a fixed project fee, a retainer. Flat price over a cost distribution that ranges 30x is a bet that your volume stays clustered near the middle of the distribution.
That bet is usually unexamined, which is a different thing from being wrong. It might be fine. Nobody has checked.
And the reason nobody has checked is structural rather than negligent. The CIO owns AI spend. The CRO owns price. Gross margin per unit of delivered work is the difference between them, and in most organizations it is not on anyone’s scorecard. The CIO is measured on total spend against budget, which creates pressure to cut. The CRO is measured on revenue and win rate, which creates pressure to sell. Neither is measured on the margin between them, so it drifts.
This is the same failure I described in why AI spending rarely reaches earnings, seen from the cost side. There the productivity gain never became profit because nothing structural changed. Here the revenue is real, and the profit leaks out through a cost line nobody connected to the price.
The four questions that establish AI unit economics
These are answerable, and answering them is usually a two-week exercise rather than a program.
1. What does it cost to deliver one unit of the thing we actually sell?
Not one API call. The completed unit the customer pays for: a resolved ticket, a processed claim, a generated report, a closed case, an onboarded account. McKinsey makes exactly this point from the CIO side, arguing that the unit of governance should be the completed business outcome rather than the token cost. That is a commercial statement, and it means the revenue side has to be in the room.
2. What does that cost look like as a distribution, not an average?
Median, seventy-fifth percentile, ninety-fifth percentile. If your data cannot produce those three numbers, that is the first finding, and it is a common one.
3. At what point in the distribution does the unit stop being profitable?
There is a token volume at which a given transaction costs more to serve than the price attributed to it. Find that threshold. It is arithmetic, not modeling.
4. What share of our volume sits past that point, and who is accountable for it?
The first half is a query. The second half is usually the harder question, and it is the one that matters. If the answer is “nobody,” you have found the actual problem, and no amount of cost optimization will fix it.
Where the answers usually lead
| What you find | What it means commercially |
|---|---|
| Cost distribution is tight | You can keep flat pricing. Most of the anxiety about AI cost does not apply to you, and you should stop spending management attention on it |
| Cost distribution is wide, but the expensive tail is a small share of volume | Flat pricing with a fair use ceiling. The ceiling exists to bound the tail, not to generate revenue, and it should be set where the tail begins rather than where it looks defensible |
| Cost distribution is wide and the expensive tail is a meaningful share | Flat pricing is subsidizing your heaviest users with your lightest ones. Move to consumption or hybrid pricing, and expect the procurement friction that comes with it |
| Cost per unit is unpredictable and correlates with customer segment | Your problem is segment selection, not pricing. Some customers are structurally unprofitable at any price you can charge them |
| Cost per unit is high but accuracy plateaus below it | You are buying tokens that do not improve the outcome. This is the one case where cost optimization alone is the right answer |
The last row deserves emphasis because it is the case most organizations assume they are in, and the Stanford finding on accuracy plateauing at intermediate cost suggests some of them genuinely are. When more tokens stop improving the answer, routing that workload to a cheaper model is free margin with no commercial consequence. That is real, and it is also the smallest of the five situations.
The cost lever is real and it is not the answer
None of this argues against AI cost management. The savings are substantial and well documented.
McKinsey reports that companies deliberate about AI consumption save 20 to 30 percent, that about a third of surveyed organizations have already achieved savings in that range, that prompt caching alone can cut repeated input-token costs by up to about 90 percent, and that organizations with high forecasting maturity save around 10 percent more than peers. Sourcing and accountability levers add 10 to 20 percent in unit cost reductions. Those are worth capturing and a competent CIO is already capturing them.
And they are targets more often than they are outcomes.
Bain surveyed 951 global companies for its Automation and AI Pathfinder Survey, published in June 2026, and asked separately what cost savings companies targeted from AI and what they actually realized. The distributions do not match.
| Annual cost savings from AI | Targeted | Realized |
|---|---|---|
| 0 to 10% | 25% | 40% |
| 11 to 20% | 37% | 29% |
| 21 to 30% | 17% | 10% |
| More than 30% | 5% | 4% |
| Not specified | 16% | 17% |
The modal company aimed at the second bucket and landed in the first. Above 20% savings, the realized share is 14% against a targeted 22%.
Two things follow, and the second is the serious one.
First, a plan whose unit economics only work at 20% cost reduction is planning on a bucket that 14% of companies reach. That is not impossible. It is a bet, and it should be labeled as one.
Second, and this is Bain’s own finding rather than an inference: 44% of companies said they plan to fund generative and agentic AI investment out of savings from prior automation programs. Out of the savings that, on the same survey’s evidence, came in below target.
Bain states the consequence plainly: the prior wave underdelivered, the savings pool is smaller than assumed, and the investment case for the current wave was sized against projections rather than actuals.
That is the same error this article describes, one level up. The gross margin problem is that nobody validates the price against the actual cost distribution. The financing problem is that nobody validates the next investment against the actual return of the last one. Both are cases of substituting the number you intended for the number you got.
Note on the two savings figures above. The McKinsey and Bain numbers are not directly comparable and should not be read as contradicting each other. McKinsey is measuring what deliberate AI consumption management achieves on AI spend, in a survey of 75 qualified respondents. Bain is measuring business-level cost savings from automation and AI programs across 951 companies. Both can be accurate. The practical reading is that optimization works when you actually do it, and that most companies do not realize what they projected, so plan against the second number and treat the first as the upside case.
But notice what those numbers are. They are percentage reductions in a cost line.
If your price is set correctly, a 25 percent cost reduction is 25 percent more margin. If your price is set against the wrong point in the cost distribution, a 25 percent cost reduction moves you from losing money on the tail to losing slightly less money on the tail, and the underlying position is unchanged.
Cost optimization improves whatever position you already hold. It does not tell you whether the position is right. Only the four questions do that, and only the commercial side can answer the fourth.
What to do in the next two weeks
Name the unit. Write down the single thing your customer pays for, in the customer’s terms. If leadership cannot agree on that sentence in one meeting, stop there, because everything downstream depends on it.
Pull the distribution. Cost per unit at the median, the seventy-fifth and the ninety-fifth percentile, for the last quarter. If the data does not exist at that granularity, the instrumentation gap is the finding, and it is fixable.
Find the break-even volume. The token count at which a unit stops paying for itself.
Count the tail. What share of last quarter’s volume sat past that line, and which customers were they.
Assign the number. One named person accountable for gross margin per unit, reporting it monthly, alongside the CIO’s spend number and the CRO’s revenue number rather than buried inside either.
That last step is the one that holds. The first four produce a slide. The fifth produces a metric, and metrics survive the quarter in a way that slides do not.
Five ways this goes wrong
- Treating it as a cost exercise and giving it to IT alone. IT can reduce spend. IT cannot reprice. A cost program without a pricing decision optimizes a position nobody validated.
- Using the average. Averages are the wrong statistic for a distribution with a long tail. The average is comfortable and the tail is where the money goes.
- Setting the fair use ceiling for revenue rather than for protection. A ceiling placed where it generates upsell revenue rather than where the unprofitable tail begins is a pricing gimmick that customers detect and resent.
- Assuming premium models are the safe default. The evidence suggests accuracy plateaus, and paying past the plateau buys nothing. This connects directly to the pricing pressure AI companies face in procurement, where a buyer who cannot predict their bill escalates the decision.
- Waiting for the data to be clean. The first pass will be approximate. An approximate distribution changes decisions. A perfect one arrives after the pricing was already set.
The shape of the problem
The comfortable framing is that AI costs are rising and need control. That framing is true, it is being written about extensively, and it puts the problem safely inside the technology function.
The uncomfortable framing is that AI moved cost of goods sold from a fixed line to a variable one with a long, unpredictable tail, and that most pricing was designed for the fixed version. That is a commercial problem, it sits between two executives who each own half of it, and it will not be solved by either of them working alone.
The companies that get this right will not be the ones with the lowest AI spend. They will be the ones that can state, from memory, what it costs them to deliver one unit of what they sell, and what happens to that number at the ninety-fifth percentile.
Talk it through
If you are selling something AI delivers and cannot currently state your gross margin per unit at the ninety-fifth percentile, that is worth an hour before the next pricing decision rather than after it.
Frequently asked questions
What are AI unit economics?
The cost to deliver one completed unit of the thing a customer pays for, measured against the price charged for it. The unit is a resolved ticket, a processed claim or a generated report, not an API call or a token. The distinguishing feature of AI unit economics is that the cost is a distribution rather than a number, because the same task can consume very different resources on different runs.
Why is AI cost per transaction so hard to predict?
Stanford’s Digital Economy Lab found that runs on the same task can differ by up to 30 times in total tokens, that human expert ratings of task difficulty align only weakly with actual token cost, and that frontier models systematically underestimate their own consumption. Agentic workflows compound this, because each step re-reads accumulated context.
Who should own AI gross margin?
One named person, reported monthly, separate from the CIO’s spend number and the CRO’s revenue number. In most organizations today it is owned by neither, which is why it drifts. The specific title matters less than the fact that a single person reports it.
Should we switch to consumption pricing?
Only after you know the distribution. If your cost distribution is tight, flat pricing is fine and switching adds procurement friction for nothing. If the distribution is wide and the expensive tail is a meaningful share of volume, flat pricing means your light users are subsidizing your heavy ones, and that is worth fixing.
Is cost optimization worth doing?
Yes, and it is not sufficient. Reported savings from deliberate AI consumption management run 20 to 30 percent, with prompt caching cutting repeated input-token costs by up to about 90 percent. Those gains improve whatever margin position you already hold. They do not tell you whether that position is correct.
How much AI cost saving do companies actually realize?
Less than they target. Bain’s Automation and AI Pathfinder Survey of 951 companies, published June 2026, found 25% targeted savings of 0 to 10% while 40% realized that range, and 37% targeted 11 to 20% while 29% realized it. Above 20% savings, 14% realized against 22% targeted. Plan against realized figures rather than targets.
Can we fund the next wave of AI investment from the savings of the last one?
Carefully, and only against actuals. Bain found 44% of companies plan to do exactly that, drawing on savings from prior automation programs, while the same survey shows those savings came in below target. Bain’s own conclusion is that the investment case for the current wave was sized against projections rather than actuals. Validate the reinvestment math against what the previous program actually returned.
Sources
- Pankaj Sachdeva and Wasim Lala, with Avinash Javaji, Kaavini Takkar and Purva Arora, “The cost of intelligence: How CIOs can manage AI demand at scale,” QuantumBlack, AI by McKinsey, July 20, 2026. Data cited from the McKinsey Enterprise AI FinOps survey, May 2026, 120 enterprise participants and 75 qualified respondents across five major industries. mckinsey.com
- Longju Bai et al., “How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks,” Stanford Digital Economy Lab, 2026. arXiv 2604.22750
- Bain and Company, “Your AI Budget Is Growing. Your Returns Aren’t. Here’s Why,” June 1, 2026. Data from the Bain Automation and AI Pathfinder Survey 2026, 951 global companies. bain.com
- Stanford Digital Economy Lab publication page. digitaleconomy.stanford.edu
APPLY THE THINKING
Turn this analysis into an accountable operating decision.
Start with the revenue, GTM, RevOps, governance, or AI constraint that matters most.
Request an AI Revenue Diagnostic