Why AI Got More Expensive
Why AI Got More Expensive
The Rule That Defined Modern Technology
For over half a century, one economic law has governed the technology industry with near-perfect reliability: everything gets cheaper.
Not marginally cheaper. Exponentially cheaper.
In 1980, storing a single gigabyte of data cost roughly $300,000 — nearly a million dollars in today’s money. By 2018, that same gigabyte cost less than two cents — a decline of more than 99.99% .
Computing power followed the same trajectory. One gigaflop of processing power cost $18.7 million in 1984. By 2017, it had fallen to $0.03 — another collapse of more than 99.99% .
The pattern repeats everywhere you look:
- Personal computers: IBM’s original PC 5150 launched in 1982 at $8,352 in inflation-adjusted dollars . By 2004, a solid entry-level desktop cost $400–500.
- Storage: A 10MB hard drive cost $7,210 (inflation-adjusted) in 1982 . Today, a 4TB drive — 400,000 times the capacity — costs $70.
- Laser printers: Hewlett-Packard’s first LaserJet cost $9,366 in today’s dollars in 1984 . Now you can pick one up for under $99.
- Solar energy: Photovoltaic costs have fallen by 90% in a single decade , with prices dropping roughly 20% every time global capacity doubled.
At the heart of it all sat Moore’s Law: the number of transistors on a chip doubling approximately every two years — a prediction that held true for more than 50 years . This relentless deflation made technology the engine of the modern economy. Entire industries were rebuilt on one safe assumption: next year’s technology will be better and cheaper.
The Three Engines of Tech Deflation
To understand why AI broke the rules, we first need to understand how the rules worked in the first place. Technology’s relentless deflation wasn’t accidental. It was driven by three powerful forces that worked in concert for over half a century.
Force #1: Moore’s Law
In 1965, Intel co-founder Gordon Moore observed that the number of transistors on a microchip was doubling approximately every two years. Remarkably, that prediction held true for more than 50 years.
What made Moore’s Law so powerful wasn’t just the density increase — it was the economic consequence. Every time transistor density doubled, the cost per transistor was roughly halved. This meant that each new generation of chips delivered dramatically more computing power at a lower unit cost. A $2,000 computer in 1995 was vastly less capable than a $500 computer in 2005, which was in turn dwarfed by a $300 smartphone in 2015.
Moore’s Law created a predictable, clockwork deflation engine for the entire computing industry.
Force #2: Economies of Scale
The second force is simpler. As manufacturing volumes increase, unit costs fall.
When the first hard drives rolled off assembly lines, they were hand-built, low-volume, and astronomically expensive. But as demand grew, factories scaled up, production processes standardized, and the cost per unit plummeted. The same pattern played out for solar panels, LED lighting, memory chips, displays, and virtually every hardware technology.
The math is brutal: a factory producing 10 million chips spreads its fixed costs differently than one producing 10,000. The difference can be several orders of magnitude in per-unit cost.
Force #3: Learning Curves (Wright’s Law)
Theodore Wright, an American engineer, first documented in 1936 that every time cumulative production of an aircraft doubled, its cost fell by a predictable percentage. This became known as Wright’s Law — or the learning curve effect.
It applies across nearly every technology: solar panels fall roughly 20% in cost for every doubling of global installed capacity. Batteries follow a similar curve. LED lighting did too. The more we produce a technology, the more we learn to produce it efficiently — and the cheaper it gets.
The self-reinforcing cycle
These three forces didn’t operate in isolation. They amplified each other:
Moore’s Law made chips cheaper → more devices became “smart” → demand exploded → economies of scale kicked in → costs fell further → more applications became viable → learning curves deepened → costs fell again.
This virtuous cycle produced the most dramatic price declines in economic history. The MIT Technology Review calculated that the cost of a transistor has fallen by a factor of roughly one million since the 1960s.
The pattern that didn’t hold
Then came AI.
In less than three years, the pattern broke. The same forces that had driven computing costs down for decades suddenly seemed to reverse.
How did a technology built on the cheapest computing power in history become more expensive than the humans it was meant to replace? The answer lies in understanding why AI, uniquely, seems to escape the gravity of Moore’s Law, economies of scale, and learning curves.
The Anomaly — AI’s Sudden Cost Surge
For decades, the evidence for tech deflation was so consistent that executives stopped questioning it. Budget for this year’s technology, and next year’s would cost less. That assumption is now colliding with an uncomfortable body of evidence — and it’s coming from the very companies building the AI revolution.
The confession from the top
In April 2026, Bryan Catanzaro, vice president of applied deep learning at Nvidia, made a statement that should be pinned to the wall of every boardroom considering an AI rollout: “For my team, the cost of compute is far beyond the costs of the employees.”
This is not a skeptic talking. This is a senior executive at the company whose chips power the entire AI industry — a company with every incentive to claim AI is cheap — admitting that the math doesn’t work. Not yet.
The academic evidence backs him up. A 2024 MIT study analyzing the technical requirements of AI models found that AI automation was economically viable in only 23% of roles where vision is a primary part of the work. In the remaining 77%, it was simply cheaper to keep humans doing the job.
Microsoft’s quiet reversal
Perhaps the most telling case study comes from a company that has bet more on AI than almost anyone. Microsoft opened access to Anthropic’s Claude Code to thousands of its developers, project managers, and designers, encouraging them to experiment. The tool became popular fast — too popular. Just six months later, Microsoft reportedly began canceling most of its direct Claude Code licenses , moving engineers toward GitHub Copilot CLI instead, because the scale of employee usage had made the costs untenable.
Let that sink in: one of the world’s largest software companies pushed AI adoption internally, and then had to pull it back because employees liked the tool too much.
Uber’s budget bonfire
Uber’s story is even starker. The company actively incentivized AI adoption through internal leaderboards ranking teams by AI tool usage. The gamification worked — spectacularly. By April, Uber had burned through its entire 2026 AI coding tools budget in just four months . As CTO Praveen Neppalli Naga admitted: “I’m back to the drawing board because the budget I thought I would need is blown away already”.
The spending paradox
And yet — the spending accelerates. Big Tech firms have announced $740 billion in AI capital expenditures so far this year, a 69% increase from 2025, according to Morgan Stanley. McKinsey estimates total AI expenditures could reach $5.2 trillion by 2030 — $1.6 trillion in data centers and $3.3 trillion in IT equipment — and as much as $7.9 trillion at an accelerated pace.
Meanwhile, the price companies pay keeps climbing. Spending management firm Tropic found that AI software fees increased by 20% to 37% over the past year .
The result is a strange economic theater: more than 118,000 tech layoffs in 2026 — already outpacing the roughly 120,000 total for all of 2025 — justified in the name of AI efficiency, even as the AI itself costs more than the workers being let go. Keith Lee, an AI and finance professor at the Swiss Institute of Artificial Intelligence’s Gordon School of Business, calls it plainly: “What we’re seeing is a short-term mismatch”.
The usage trap
Here is where the anomaly becomes a paradox. Having spent billions on AI, companies are now pressuring employees to use as much of it as possible. Amazon is pushing staff to “toxenmaxx” — consume as many AI tokens as possible . A Meta employee built an internal leaderboard, wryly named “Claudeonomics,” to rank workers by AI usage. Nvidia CEO Jensen Huang has predicted that 100 AI agents will one day work alongside every one of his employees.
But under token-based pricing, usage is the bill. Every prompt, every agent task, every generated line of code adds to the meter. Companies are simultaneously discovering that their AI costs scale with success — the more value employees get from the tools, the larger the invoice at the end of the month.
This is the anomaly in full: a technology built on the cheapest computing substrate in human history, whose costs to the enterprise are rising, not falling. In the next section, we’ll examine three hypotheses for why AI escapes the economic gravity that pulled every other technology toward zero — and why the answer matters enormously for how you budget the next five years.
Why Does AI Escape the Economic Rule?
Every previous technology that deflated — transistors, storage, solar panels, PCs — shared one thing in common: none of them escaped the three forces we just described. So why is AI different? Here are three hypotheses, each explaining a different part of the anomaly. The honest answer is probably a combination of all three.
Hypothesis 1: The Impossible Quest for AGI
The modern AI industry was built on two assumptions that have quietly collapsed.
The first was that this generation of AI would achieve Artificial General Intelligence — a machine capable of matching or exceeding human cognition across all domains. The second was that whoever reached AGI first would enjoy a winner-take-all monopoly, justifying almost any level of spending to get there.
Neither assumption holds today. Most leading AI companies no longer mention AGI as a strategy. They are competing on benchmarks, user adoption, and compute scale. And far from converging on a single winner, the field has settled into a crowded oligopoly — OpenAI, Anthropic, Google, Meta, Microsoft, and a growing wave of open-source challengers — with no monopoly in sight.
But the spending logic born in the AGI era hasn’t died with the AGI dream. The industry remains locked in an arms race of scale: bigger models, bigger clusters, bigger data centers. The result is an upward spiral of infrastructure investment that has little to do with what customers actually need — and everything to do with competitive positioning. Big Tech has announced $740 billion in AI capital expenditures so far this year, a 69% increase from 2025, according to Morgan Stanley . McKinsey projects total AI expenditures could reach $5.2 trillion by 2030 — and as much as $7.9 trillion at an accelerated pace .
When an industry spends like it’s racing for a monopoly that will never exist, somebody has to pay for it. That somebody is the customer.
Hypothesis 2: The Token Paradox — Jevons Paradox for AI
In 1865, economist William Stanley Jevons observed something counterintuitive: as steam engines became more efficient and coal became cheaper to use, Britain’s total coal consumption didn’t fall — it exploded. Efficiency reduced the cost per unit of work, but it made so many new uses economically viable that total demand grew faster than efficiency gains.
AI is now living its own Jevons Paradox.
On one side, the unit economics are improving exactly as Moore’s Law would predict. Gartner forecasts that by 2030, inference on a one-trillion-parameter LLM will cost AI firms nearly 90% less than it did in 2025. On the other side, consumption is growing even faster: Goldman Sachs projects that agentic AI could drive a 24-fold increase in token consumption by 2030, reaching a staggering 120 quadrillion tokens per month .
The math is brutal in its simplicity: a 90% price cut sounds dramatic until you multiply it against 24× the volume. Your per-token cost falls by an order of magnitude; your total bill still goes up.
And the industry is actively engineering this outcome. Nvidia CEO Jensen Huang has predicted that 100 AI agents will one day work alongside every one of his employees . Each agent is a token-consuming machine, running tasks in parallel, around the clock, at machine speed.
Gartner’s warning to product leaders deserves to be read slowly: “Chief Product Officers should not confuse the deflation of commodity tokens with the democratization of frontier reasoning”. Cheaper tokens will not produce cheaper AI, because agentic models consume far more tokens per task, consumption growth outpaces unit-cost decline, and providers won’t fully pass savings through to customers.
This is the crucial difference from every deflationary technology before it. Hard drives got cheaper, and nobody suddenly needed 10,000× more storage per task. AI tokens get cheaper, and the industry responds by building agents that consume them by the quadrillion.
Hypothesis 3: The Pricing Power of AI Vendors
The third hypothesis is the least technical and the most uncomfortable: AI is expensive because the companies selling it can make it expensive.
Moore’s Law operated in fiercely competitive commodity markets. Memory chips, storage, CPUs — dozens of manufacturers, brutal price competition, near-zero switching costs. When your cost per transistor fell, you passed the savings on or your competitor did.
The AI market looks nothing like that. The model layer is controlled by a handful of players — OpenAI, Anthropic, Google, Microsoft — and the hardware layer is dominated by a single company. When supply is concentrated and demand is exploding, prices don’t follow cost curves. They follow what the market will bear.
The evidence is already visible in your software budget. Fees for AI software increased by 20% to 37% over the past year , according to spending management firm Tropic — at a time when the underlying compute was supposedly getting cheaper.
Worse, the current pricing models are broken in both directions. Flat subscriptions fail to cover operating costs for heavy users, meaning providers are losing money on their most enthusiastic customers. The predictable consequence, as professor Keith Lee notes, is a coming shift from flat subscriptions to usage-based pricing [1] — which transfers all the consumption risk directly onto you, the buyer.
Small wonder, then, that some firms are beginning to reevaluate AI “not as a clear cost-saving substitute for labor, but as a complementary tool — at least until the cost structure stabilizes”.
The synthesis: An industry overspending on an arms race (Hypothesis 1), built on a consumption model where cheaper units mean bigger bills (Hypothesis 2), sold by vendors with the market power to keep prices high (Hypothesis 3). Any one of these would be enough to break the deflationary rule. AI has all three.
But AI is not the first technology to defy gravity. In the next section, we’ll examine three others that broke the same rule — and what their trajectories tell us about where AI costs are heading.
Three Other Technologies That Broke the Same Rule
AI is not the first technology to defy the deflationary rule. Three others have done it before — and their stories illuminate both why AI is expensive today and where its costs are heading tomorrow.
Exception #1: Enterprise SaaS
Software-as-a-Service was supposed to be the ultimate deflationary product: infinitely copyable, near-zero marginal cost to serve one more customer. And yet SaaS prices are rising roughly 11.4% per year — more than four times the average inflation rate of G7 countries . Software inflation is running at more than double the US CPI, and some vendors — particularly those acquired by private equity firms — have imposed increases as high as 900%.
Why? Because SaaS doesn’t sell a commodity. It sells a dependency. Once your workflows, your data, and your team’s daily habits live inside a platform, switching costs become enormous. Vendors know this, and they price accordingly.
The AI parallel: Model providers hold exactly the same leverage. Once your company builds its processes around a specific vendor’s models, prompts, and agents, you are no longer a customer — you are a captive. AI software fees already rose 20% to 37% in a single year . The subscription price you signed today is not the price you’ll pay at renewal.
Exception #2: Advanced Chip Manufacturing
The semiconductor industry is the birthplace of Moore’s Law — and, ironically, the place where its economic promise broke first.
A mature 28nm wafer from TSMC costs around $3,000. A cutting-edge 3nm wafer costs roughly $19,500 — about 6.5 times more . Each EUV lithography machine required to print these nodes costs over $350 million. An advanced AI accelerator like NVIDIA’s H100 carries an estimated manufacturing cost of $3,320, versus $3–10 for a mature-node chip — a thousandfold increase.
For decades, the cost per transistor kept falling. But at the leading edge, the physics became so difficult, and the tools so expensive, that the absolute cost of frontier chips started climbing — even as older nodes stayed cheap.
The AI parallel: AI is chasing the frontier in precisely the same way. The industry’s biggest spenders aren’t buying commodity compute; they’re buying frontier compute, where every increment of capability costs a multiple of the last. Big Tech’s $740 billion capex surge — a 69% jump in a single year — is frontier pricing, not commodity pricing.
Exception #3: College Tuition
Not a hardware technology, but perhaps the most instructive exception of all. US college tuition has risen more than 32-fold since 1963 — up 229.8% even after adjusting for inflation — climbing an average of 5.8% annually since 1983, faster than almost any other household expense.
Economists explain it through a combination of dynamics: the Bennett Hypothesis (the more money available — in that case, student aid — the higher prices climb), oligopolistic competition (a handful of prestigious institutions sets the price ceiling for everyone else), and a product perceived as indispensable, where demand barely responds to price at all. Nobody wants to buy the cheaper, lesser-known degree.
The AI parallel: The same psychology now drives corporate AI adoption. Companies aren’t buying AI because the ROI is proven — Federal Reserve data shows only about 18% of companies had adopted AI tools by the end of 2025 . They’re buying it because they fear being the only ones who didn’t. And just as easy student aid inflated tuition, easy capital — $740 billion and counting — inflates the entire AI cost structure.
The Common Anatomy
Lay the exceptions side by side, and a pattern emerges:
| Exception | Constrained Supply | Inelastic or Exploding Demand | Vendor Pricing Power |
|---|---|---|---|
| Enterprise SaaS | Few dominant platforms per category | Switching costs lock customers in | ~11.4% annual price increases |
| Leading-edge chips | One dominant foundry; one EUV supplier | AI boom strains frontier capacity | 6.5× wafer cost premium |
| College tuition | Elite institutions limited by prestige | Degree seen as non-negotiable | 5.8% annual increases for 40 years |
| AI | Handful of model providers; one dominant chipmaker | FOMO-driven adoption; token consumption up 24× by 2030 | Fees up 20–37% in one year |
Every technology that escaped the deflationary rule shares this anatomy: constrained supply, demand that doesn’t respond to price, and sellers who know it. AI checks all three boxes — which tells us the current cost surge is structural, not a temporary accident.
But structural does not mean permanent. SaaS monopolies get disrupted. Chip nodes mature and commoditize. Even tuition inflation eventually slowed. The question for your business is not whether AI economics will normalize — it’s when, and what you should do in the meantime. That’s the subject of the next section.
A Cautious Prediction — What Happens Next?
Predicting the future of AI pricing is a fool’s errand. But predicting its shape is not. The evidence we’ve assembled — from AI’s own economics and from the three technologies that broke the same rule before it — points to five developments that CEOs should treat as planning assumptions, not possibilities.
1. The hardware layer will deflate — genuinely
The physics of compute have not changed. Gartner projects that by 2030, performing inference on a one-trillion-parameter LLM will cost nearly 90% less than it did in 2025. Specialized AI chips, maturing data center designs, and improving model architectures will do to AI compute what Moore’s Law did to general-purpose chips. If someone tells you the unit cost of AI will keep rising forever, the data says otherwise.
2. The software layer will not follow
Here is the catch, and Gartner states it bluntly: cheaper tokens will not translate into cheaper enterprise AI, because agentic models consume far more tokens per task, consumption growth outpaces falling unit costs, and — critically — providers won’t fully pass the savings through to customers. Meanwhile, the broken flat-subscription model that loses money on heavy users will be replaced. Professor Keith Lee predicts a shift from flat subscriptions to usage-based pricing — a change that moves all consumption risk from the vendor’s balance sheet onto yours.
3. The Jevons Paradox will dominate your budget
Do not mistake falling unit prices for falling bills. Goldman Sachs forecasts a 24-fold increase in token consumption by 2030 — reaching 120 quadrillion tokens per month as AI agents proliferate. A 90% price cut multiplied by 24× the volume is still a bill that more than doubles. For the next three to five years, total enterprise AI spending will rise even as per-token costs collapse. Plan your budgets for growth, not deflation.
4. A consolidation is coming
The current spending level is not a steady state — it’s a land grab. Big Tech has announced $740 billion in AI capital expenditures this year alone, with McKinsey projecting up to $5.2 trillion in cumulative AI expenditures by 2030. Arms races of this magnitude do not end with everyone still standing. Some model providers will be acquired, some will pivot, and some will simply run out of capital. The survivors will be those who can deliver predictable, sustainable pricing — and every one of their stranded customers will face a painful migration. This is not speculation; it is the documented fate of every previous technology gold rush.
5. The real tipping point will be reliability, not price
The most important prediction is also the most easily missed. The moment AI becomes a genuine labor substitute won’t be marked by a price crossing — it will be marked by a trust crossing. As Keith Lee puts it: “It’s not just about AI becoming cheaper than humans. It’s about becoming both cheaper and more predictable at scale”. That means fewer hallucinations, reduced need for human oversight, and clean integration into existing infrastructure. Adoption is already accelerating but the firms winning today are not the ones with the most agents. They’re the ones with the most dependable ones.
The bottom line of the forecast: hardware costs fall, vendor prices stay sticky, consumption explodes, the vendor landscape consolidates, and value concentrates in reliable deployment rather than raw capability. Each of these has a direct implication for how you should be spending — and protecting — your AI budget right now. That is the subject of our final section: concrete advice for CEOs navigating the surge.
The CEO’s Playbook — Six Moves for Navigating the AI Cost Surge
If you’ve read this far, you understand the landscape: AI costs are rising even as unit prices fall, vendors hold unprecedented pricing power, and your peers are burning through budgets at record speed. The question is no longer whether to adopt AI — the adoption train has left the station, with 18% of companies using AI tools by the end of 2025 and the adoption rate growing 68% in a single quarter . The question is how to adopt it without becoming the next cautionary tale.
Here are six moves we recommend to CEOs of mid-to-large companies.
1. Don’t “Toxenmaxx” Blindly
Amazon is pushing employees to consume as many AI tokens as possible. Meta built a leaderboard — wryly named “Claudeonomics” — to rank workers by AI usage. Uber gamified adoption with internal leaderboards and burned through its entire 2026 AI coding tools budget by April .
The lesson is not that these companies failed at adoption. It’s that they optimized for the wrong metric. Under token-based pricing, usage is the bill. Rewarding token volume is like rewarding employees for maximizing their electricity consumption — you’ll succeed, and you’ll regret it.
Your move: Measure ROI per token, not volume of tokens. Every AI initiative should report a simple ratio: business value generated divided by tokens consumed. Teams that deliver outcomes with lean prompts and right-sized models should be your heroes — not the ones topping the consumption leaderboard.
2. Separate the Hardware Curve from the Vendor Curve
The most dangerous budgeting error you can make is assuming that because AI compute is getting cheaper, your AI bill will too. Gartner projects inference costs will fall nearly 90% by 2030 — and simultaneously warns that cheaper tokens won’t translate into cheaper enterprise AI , because agentic workloads consume vastly more tokens and providers won’t fully pass savings through.
Your move: Treat vendor pricing as a negotiation, not a law of nature. Push for multi-year contracts with explicit pricing caps and defined escalation clauses. Expect vendors to migrate from flat subscriptions — which are losing them money on heavy users — to usage-based pricing, and model what that shift means for your budget before signing.
3. Build Cost Governance Before You Scale
Microsoft opened Claude Code to thousands of employees, encouraged experimentation — and reversed course within six months when usage outgrew the budget, moving engineers to GitHub Copilot CLI instead. If a company with Microsoft’s resources and sophistication can be caught off-guard, so can you.
Your move: Deploy usage budgets, approval workflows, and real-time cost dashboards before broad rollouts, not after the first budget blowout. Set per-team and per-project token caps. Make AI spend as visible and accountable as headcount spend. The companies that govern AI consumption early will scale it sustainably; the ones that don’t will be forced into chaotic, morale-damaging reversals.
4. Architect for Model Portability
Vendor lock-in is the silent multiplier of AI costs. If your infrastructure can only run one vendor’s models, you have no defense when that vendor raises prices or retires the model you depend on. Microsoft’s own pivot between Claude Code and Copilot CLI shows that even the largest players shuffle their model stack — and you want the same freedom.
Your move: Build an abstraction layer between your business logic and the underlying models. Require that prompts, agents, and workflows can be migrated across providers in weeks, not quarters. When a cheaper model matches your quality bar — and they will, increasingly often — you should be able to switch in days.
5. Budget AI as Augmentation, Not Replacement
The most common budgeting mistake is assuming AI will pay for itself by replacing headcount. The data says otherwise — for now. MIT found AI automation economically viable in only 23% of vision-based roles. Even Nvidia’s own VP admits compute costs exceed employee costs on his team. Meanwhile, the tech sector has already cut more than 118,000 jobs in 2026 in part to fund AI investments — layoffs justified by savings that haven’t materialized.
Your move: Follow professor Keith Lee’s framing: treat AI as “a complementary tool — at least until the cost structure stabilizes” . Budget it as a productivity investment with its own ROI model, measured in output quality and speed — not as a headcount substitution line. When the cost crossover does come, you’ll be positioned to capture it. If you cut first and wait for AI to fill the gap, you’re betting the company on a price curve you don’t control.
6. Invest in Integration, Not Just Licenses
Here’s the insight that ties everything together. The companies bleeding money on AI are the ones consuming it raw — renting generic tools on consumption pricing, with usage spiraling and value diffuse. The companies getting durable returns are the ones embedding AI into specific, high-value workflows, connected to their own data and systems, where every token spent maps to a business outcome.
Your move: Prioritize depth over breadth. Pick a small number of workflows where AI demonstrably outperforms the status quo, integrate deeply with your legacy systems, and instrument everything. A focused, well-integrated AI deployment with hard ROI data is worth more — and costs less — than an enterprise-wide license rollout with a leaderboard attached.
The Bottom Line
The AI cost surge is real, but it is not a reason to retreat. It is a reason to be deliberate.
The pattern from every technology that came before is clear: costs spike during the land-grab phase, then normalize as infrastructure matures, competition consolidates, and vendors learn to price sustainably. Inference costs will fall 90%. Model capability will keep rising. The gap between AI’s cost and its value will close — the only question is whether your company will still have the budget, the governance, and the organizational trust to capitalize when it does.
As Keith Lee reminds us: “It’s not just about AI becoming cheaper than humans. It’s about becoming both cheaper and more predictable at scale”. Predictability is exactly what you should demand — from your vendors, from your governance, and from your own AI strategy.
The winners of this phase won’t be the companies that adopted AI fastest. They’ll be the ones that adopted it most intelligently — with clear-eyed budgets, ruthless measurement, and architectures that keep them free to move as the economics shift.
FAQ: The Executive’s Guide to Navigating the AI Cost Surge
Q1: Why is AI getting more expensive when every other technology in history has gotten cheaper? A: For over 50 years, technology followed a predictable deflationary rule driven by three forces: Moore’s Law (transistor density doubling every two years), economies of scale (falling unit costs as production scales), and learning curves (falling costs as cumulative production doubles). These forces made transistors, storage, and computing power over 99.99% cheaper. AI breaks this pattern because it is subject to three countervailing forces: an industry locked in an arms race for frontier models, a Jevons Paradox where cheaper tokens drive exponentially more consumption, and vendor pricing power concentrated among a handful of dominant providers. The result is that AI software fees rose 20–37% in a single year, even as underlying compute costs fell — a structural anomaly, not a temporary one.
Q2: What is the Jevons Paradox, and why does it mean my AI bill will keep growing even as prices fall? A: The Jevons Paradox, first observed in 1865, states that as a resource becomes more efficient to use, total consumption doesn’t fall — it explodes, because efficiency makes new use cases economically viable. AI is living this paradox today. Gartner projects inference costs will fall nearly 90% by 2030, but Goldman Sachs forecasts a 24-fold increase in token consumption over the same period, driven by the proliferation of AI agents. The math is simple: a 90% price cut multiplied by 24× the volume means your total bill still more than doubles. Nvidia CEO Jensen Huang has predicted 100 AI agents will work alongside every employee — each one a token-consuming machine running around the clock. Cheaper tokens do not mean cheaper AI when consumption grows faster than unit costs decline.
Q3: What does the Nvidia VP’s comment about compute costs exceeding employee costs tell us about the economics of AI? A: In April 2026, Bryan Catanzaro, Nvidia’s VP of applied deep learning, stated: “For my team, the cost of compute is far beyond the costs of the employees.” This is a landmark admission from the company whose chips power the entire AI industry — a company with every incentive to claim AI is cheap. It confirms that at the current cost structure, AI is not yet a viable labor substitute for most tasks. An MIT study from 2024 reinforces this, finding AI automation economically viable in only 23% of vision-based roles — in the remaining 77%, it was simply cheaper to keep humans. For executives, this means budgeting AI as augmentation (a productivity investment with its own ROI model), not as a headcount replacement strategy.
Q4: What happened with Microsoft’s Claude Code rollout, and what should I learn from it? A: Microsoft opened access to Anthropic’s Claude Code to thousands of developers, project managers, and designers, encouraging broad experimentation. The tool became popular fast — too popular. Just six months later, Microsoft began canceling most of its direct Claude Code licenses, moving engineers toward GitHub Copilot CLI instead, because employee usage had made costs untenable. The lesson: one of the world’s largest software companies pushed AI adoption internally and then had to pull it back because employees liked the tool too much. This underscores the critical need for cost governance before broad rollout — setting usage budgets, approval workflows, and real-time cost dashboards before the first budget blowout, not after.
Q5: Uber burned through its entire 2026 AI budget in four months. How did that happen? A: Uber actively incentivized AI adoption through internal leaderboards ranking teams by AI tool usage. The gamification worked spectacularly — and by April, Uber had burned through its entire 2026 AI coding tools budget. As CTO Praveen Neppalli Naga admitted: “I’m back to the drawing board because the budget I thought I would need is blown away already.” The root cause is what we call the “usage trap”: under token-based pricing, usage is the bill. Rewarding token volume is like rewarding employees for maximizing electricity consumption. The correct metric is ROI per token — business value generated divided by tokens consumed. Teams that deliver outcomes with lean prompts and right-sized models should be the heroes, not the ones topping consumption leaderboards.
Q6: How much is Big Tech actually spending on AI, and what does that mean for my company’s costs? A: Big Tech has announced $740 billion in AI capital expenditures in 2026 alone — a 69% increase from 2025, according to Morgan Stanley. McKinsey projects total AI expenditures could reach $5.2 trillion by 2030 ($1.6 trillion in data centers and $3.3 trillion in IT equipment), and as much as $7.9 trillion at an accelerated pace. This is not a market that will converge on a single winner — the field has settled into a crowded oligopoly of OpenAI, Anthropic, Google, Meta, Microsoft, and open-source challengers. When an industry spends like it’s racing for a monopoly that will never exist, the customer pays. This spending inflates the entire cost structure, and until consolidation occurs, enterprise AI pricing will remain elevated.
Q7: What are the three hypotheses for why AI escapes the economic gravity that pulled every other technology toward zero? A: Three hypotheses explain the anomaly, and the honest answer is a combination of all three:
-
The Impossible Quest for AGI: The industry built spending logic around achieving Artificial General Intelligence and a winner-take-all monopoly. Neither assumption holds today, but the $740 billion capex arms race continues, disconnected from what customers actually need.
-
The Token Paradox (Jevons Paradox): Unit costs are falling (90% projected by 2030), but token consumption is growing 24× faster. The per-token price drops; the total bill grows. Agentic models consume vastly more tokens per task, and providers won’t fully pass savings through.
-
Vendor Pricing Power: Unlike the brutally competitive commodity markets that drove Moore’s Law, the AI market is concentrated among a handful of model providers and one dominant chipmaker. When supply is concentrated and demand is exploding, prices follow what the market will bear — not cost curves. Evidence: AI software fees rose 20–37% in a single year.
Q8: What three other technologies broke the deflationary rule, and what do they teach us about AI’s trajectory? A: Three technologies escaped the deflationary rule before AI, and their trajectories are instructive:
-
Enterprise SaaS: Software with near-zero marginal cost of distribution should be deflationary, yet SaaS prices rise ~11.4% annually. Why? Once your workflows and data live inside a platform, switching costs lock you in. Vendors know this. AI model providers hold exactly the same leverage.
-
Leading-edge Chip Manufacturing: A mature 28nm wafer costs ~$3,000; a cutting-edge 3nm wafer costs ~$19,500 — 6.5× more. The cost per transistor fell, but the absolute cost of frontier chips climbed. AI is chasing the same frontier, where every increment of capability costs a multiple of the last.
-
College Tuition: US college tuition has risen 229.8% above inflation since 1963. The drivers — FOMO, oligopolistic pricing, and a product seen as indispensable — now drive corporate AI adoption. Only 18% of companies had adopted AI tools by end of 2025, yet the spending is relentless.
All three share a common anatomy: constrained supply, inelastic demand, and vendor pricing power. AI checks all three boxes, confirming the cost surge is structural.
Q9: When will AI costs actually start coming down, and what’s the realistic timeline for normalization? A: The evidence points to five developments that CEOs should treat as planning assumptions:
-
Hardware costs will deflate genuinely — Gartner projects inference costs will fall ~90% by 2030. The physics of compute has not changed.
-
Software costs will not follow — cheaper tokens won’t translate to cheaper enterprise AI because agentic workloads consume more tokens and providers won’t fully pass savings through.
-
The Jevons Paradox will dominate budgets — a 90% price cut × 24× volume means bills more than double. Plan for growth, not deflation, for the next 3–5 years.
-
A consolidation is coming — $740 billion arms races don’t end with everyone standing. Some providers will be acquired, some will fail. The survivors will offer predictable pricing; every stranded customer will face painful migration.
-
The real tipping point is reliability, not price — AI becomes a genuine labor substitute not when it’s cheaper, but when it’s both cheaper and more predictable at scale. The winning firms are the ones with the most dependable agents, not the most agents.
Q10: How should I budget for AI over the next 3–5 years? A: Budget for growth, not deflation. The single most dangerous budgeting error is assuming that because AI compute is getting cheaper, your AI bill will too. It won’t — for three reasons: agentic models consume more tokens per task, consumption growth outpaces unit-cost decline, and providers won’t fully pass savings through.
Follow professor Keith Lee’s framing: treat AI as “a complementary tool — at least until the cost structure stabilizes.” Budget it as a productivity investment with its own ROI model, measured in output quality and speed, not as a headcount substitution line. The tech sector has already cut more than 118,000 jobs in 2026 in part to fund AI investments — layoffs justified by savings that haven’t materialized. Don’t bet the company on a price curve you don’t control.
Q11: What’s the single most important metric I should track for AI initiatives? A: ROI per token — business value generated divided by tokens consumed. Every AI initiative should report this simple ratio. Under token-based pricing, usage is the bill. Teams that deliver outcomes with lean prompts, right-sized models, and efficient architectures should be your heroes — not the ones topping consumption leaderboards. Make AI spend as visible and accountable as headcount spend. Companies that govern AI consumption early will scale it sustainably; those that don’t will be forced into chaotic, morale-damaging reversals — as Microsoft, Uber, and others have already experienced.
Q12: How do I protect my company from vendor lock-in and rising AI software prices? A: Build an abstraction layer between your business logic and the underlying models. Require that prompts, agents, and workflows can be migrated across providers in weeks, not quarters. This is called “model portability,” and it’s your single best defense against vendor pricing power. When a cheaper or better model matches your quality bar — and they will, increasingly often — you should be able to switch in days.
Also, push for multi-year contracts with explicit pricing caps and defined escalation clauses. Expect vendors to migrate from flat subscriptions (which lose money on heavy users) to usage-based pricing, and model what that shift means for your budget before signing. Microsoft’s own pivot between Claude Code and Copilot CLI shows that even the largest players shuffle their model stack — you want the same freedom.
Q13: Should I be worried about the 118,000+ tech layoffs in 2026 being justified by AI efficiency? A: Yes, but not for the reasons you might think. The tech sector has cut more than 118,000 jobs in 2026 — already outpacing 2025’s total of roughly 120,000 — many justified in the name of AI efficiency. Yet the AI itself often costs more than the workers being let go. As professor Keith Lee calls it: “a short-term mismatch.”
The risk is that companies cut headcount expecting AI to fill the gap, only to discover the AI costs more than the people they replaced — and delivers less value. The correct approach is to budget AI as augmentation, not replacement. When the cost crossover does come (and it likely will), you’ll be positioned to capture it. If you cut first and wait for AI to fill the gap, you’re betting the company on a price curve you don’t control.
Q14: What’s the difference between companies that are getting value from AI and those that are bleeding money on it? A: The companies bleeding money are consuming AI raw — renting generic tools on consumption pricing, with usage spiraling and value diffuse. They optimize for token volume, gamify adoption, and measure nothing. The companies getting durable returns are embedding AI into specific, high-value workflows, connected to their own data and legacy systems, where every token spent maps to a business outcome. They prioritize depth over breadth: a small number of deeply integrated workflows with hard ROI data, rather than enterprise-wide license rollouts with leaderboards attached. The winners of this phase won’t be the companies that adopted AI fastest. They’ll be the ones that adopted it most intelligently — with clear-eyed budgets, ruthless measurement, and architectures that keep them free to move as the economics shift.
Q15: What is the CEO’s six-move playbook for navigating the AI cost surge? A: Based on the evidence from the article, here are the six moves:
-
Don’t “Toxenmaxx” Blindly — Measure ROI per token, not volume of tokens. Rewarding token consumption is like rewarding employees for maximizing electricity usage.
-
Separate the Hardware Curve from the Vendor Curve — AI compute is getting cheaper; your AI bill isn’t. Treat vendor pricing as a negotiation, not a law of nature. Push for multi-year contracts with pricing caps.
-
Build Cost Governance Before You Scale — Deploy usage budgets, approval workflows, and real-time cost dashboards before broad rollouts, not after the first budget blowout. Make AI spend as visible as headcount spend.
-
Architect for Model Portability — Build an abstraction layer so prompts, agents, and workflows can migrate across providers in weeks. When a cheaper model matches your quality bar, you should be able to switch in days.
-
Budget AI as Augmentation, Not Replacement — Treat AI as a complementary tool with its own ROI model until the cost structure stabilizes. Don’t cut headcount expecting AI to fill the gap.
-
Invest in Integration, Not Just Licenses — Prioritize depth over breadth. Pick a small number of workflows where AI demonstrably outperforms the status quo, integrate deeply with legacy systems, and instrument everything. A focused, well-integrated deployment with hard ROI data costs less and delivers more than an enterprise-wide license rollout.
The bottom line: the AI cost surge is real, but it is not a reason to retreat. It is a reason to be deliberate. The winners will be the companies that adopt AI most intelligently — not fastest.
We are Here to Empower
At System in Motion, we are on a mission to empower as many knowledge workers as possible. To start or continue your GenAI journey.
You should also read
The Regional AI Paradox: Build a Bridge
Article 17 minutes readLet's start and accelerate your digitalization
One step at a time, we can start your AI journey today, by building the foundation of your future performance.
Book a Training