There is an AI bubble forming. That statement usually triggers one of two reactions.
- The first is that AI is obviously transformative, demand is exploding, and therefore there cannot be a bubble.
- The second is that AI is overhyped, enterprises will eventually realize it and the entire market will collapse.
I think both arguments miss what is actually happening.
AI works. Enterprise adoption is growing. Inference demand is accelerating. AI is becoming embedded into cloud infrastructure, software development, cybersecurity, advertising, productivity applications and business processes. But none of that prevents a financial bubble from forming around it.
The potential bubble isn’t AI. It is the amount of capital being deployed ahead of proven economic returns from AI. And the numbers are becoming enormous.
Follow the Capital
The largest technology companies are conducting one of the biggest infrastructure buildouts in technology history.
Amazon, Microsoft, Alphabet Inc. and Meta are spending extraordinary amounts on data centers, GPUs, custom accelerators, CPUs, memory, networking, cooling, power systems and the physical infrastructure required to connect all of it.
The GPU gets the attention.
The system gets the bill.
And the bill increasingly extends beyond computing into power generation, transmission, substations, transformers, cooling and land. The scale is easier to understand when the major platforms are viewed together.
One caveat is critical. These numbers should not be added together and labeled “AI capex.”
Hyperscale infrastructure is fungible. Microsoft infrastructure supports Azure, Microsoft 365, GitHub and other products. Google’s supports Cloud, Search, advertising, Gemini and DeepMind. Meta uses AI throughout its advertising and recommendation systems. Amazon deploys infrastructure across AWS and its broader businesses.
The correct question therefore isn’t:
How much AI revenue exists relative to AI capex?
It is: How much incremental economic return is being generated by the capital being deployed because of AI? That is much harder to answer. But it is ultimately the number that matters.
Demand Is Not the Problem Today
- AWS illustrates why dismissing the AI infrastructure boom as irrational would be a mistake.
- AWS grew 36.7% in Q2 2026, its fastest growth in 18 quarters, reaching a $169 billion annualized revenue run rate.
- More importantly, Amazon says AWS’s AI business has exceeded a $25 billion annual revenue run rate and is growing at triple-digit percentages.
- Its chips business has also crossed a $25 billion annual run rate. That is real monetization.
- NVIDIA provides even more dramatic evidence. Its Q1 FY27 Data Center revenue reached $75.2 billion, up 92% year over year.
- The market isn’t spending hundreds of billions because nobody is using AI. Quite the opposite.
The uncomfortable question is whether today’s growth rates can eventually support the amount of capital chasing that demand.
Scarcity Is Delaying the Moment of Truth
- One of the most important dynamics in AI infrastructure today is scarcity.
- When GPU capacity is constrained, almost every available GPU can find a workload.
- When HBM is constrained, pricing rises.
- When power is constrained, customers reserve capacity years before facilities become operational.
- When transformers and switchgear have long lead times, the bottleneck shifts from silicon to electrical infrastructure.
All of this creates an environment where constrained supply can look like unlimited demand. But those are different things.
Today’s market tells us that demand exceeds available supply.
- It does not tell us what equilibrium demand looks like when supply becomes broadly available.
- Memory may provide the clearest evidence. Micron’s fiscal Q3 2026 revenue reached approximately $41.5 billion, compared with $9.3 billion a year earlier. Its GAAP gross margin reached 84.6%. Its Cloud Memory business generated an 83% gross margin, while its Core Data Center business reached 87%.
- Micron then guided the following quarter to roughly $50 billion of revenue at approximately 86% gross margin. Those are extraordinary numbers. They demonstrate how strategically important memory has become to AI infrastructure. They also demonstrate the economics of scarcity.
- The question isn’t whether memory demand is real. It clearly is.
- The question is what those economics look like when supply catches demand.
- That same question eventually applies to GPUs.
The Price of Intelligence Is Falling
- This is where the AI infrastructure equation becomes particularly interesting.
- While the capital cost of building AI infrastructure is exploding, the unit cost of consuming AI intelligence is moving in the opposite direction.
- Models become more efficient.
- Inference software improves.
- Quantization reduces computational requirements.
- Distillation produces smaller models.
- Caching eliminates redundant computation.
- Speculative decoding increases throughput.
- Custom silicon improves economics.
- Enterprises increasingly route workloads across different models.
And each new generation of hardware delivers substantially better performance per dollar and performance per watt. NVIDIA itself is telling us where this is going. The company says its Rubin platform can deliver up to a 10x reduction in inference token cost compared with Blackwell. Think about the implications. That is fantastic for AI adoption. But it creates a brutal economic treadmill for infrastructure owners.
The industry is spending unprecedented amounts of money to manufacture AI computation while simultaneously engineering down the unit price of that computation. For the economics to work, consumption must increase faster than the unit cost of intelligence declines. That is entirely possible. But it is not guaranteed.
GPU Count Is the Wrong Metric
The industry loves counting infrastructure.
- GPUs.
- FLOPs.
- Megawatts.
- Gigawatts.
- Tokens.
- Clusters.
None of those measures tells us whether the infrastructure generates an acceptable economic return. A GPU can be installed and powered on without being economically productive. Training systems lose efficiency through communication overhead, memory bottlenecks, synchronization, checkpointing and data movement. Inference introduces batching, latency requirements, KV-cache management, memory bandwidth and workload variability.
Agents make the problem even more interesting. An agent may invoke a model, query a database, wait for an API, execute code, invoke another model and then wait for an external system.
The application can be working while the GPU is waiting. This is why theoretical FLOPs are an incomplete economic metric. What matters isn’t how much compute exists. It is how much useful business output that compute produces.
The metric I would increasingly watch is economic output per GPU-hour. That moves AI infrastructure management from capacity planning toward FinOps. And I believe that transition is coming much faster than many infrastructure providers expect.
Then Comes the Second Bill: Depreciation
- Capex is only the first bill.
- Depreciation is the second.
- Every AI server installed today becomes an expense flowing through tomorrow’s income statement.
- And AI accelerators have a particularly interesting economic characteristic.
- They can become economically obsolete long before they become physically obsolete.
Imagine a Blackwell cluster that continues functioning perfectly five years from now. If Rubin or its successors can produce substantially more inference per watt and dramatically lower cost per token, that older cluster still works. But economically, it must compete against newer infrastructure with radically better unit economics. That creates an enormous question around useful asset life.
We may be building one of the largest pools of rapidly depreciating technology assets ever created at exactly the moment when AI compute efficiency is improving at extraordinary speed. That doesn’t mean the assets become worthless. It means they have to be utilized aggressively.
The AI Economic Breakeven Problem
This is where I think the AI bubble debate needs to become more quantitative. Suppose the industry deploys $500 billion of incremental AI-related infrastructure capital. This is a thought experiment, not a claim that $500 billion of disclosed hyperscaler capex is purely AI. What return must that capital produce?
This is deliberately simplified. It doesn’t account fully for taxes, financing structure, replacement capex, residual asset values or differences between accounting and economic depreciation.
But that is exactly why the table is useful. It shows the magnitude of the economic problem.
- At a 15% target return, $500 billion of capital must eventually produce roughly $75 billion of annual economic return.
- At a 30% operating margin, that implies roughly $250 billion of incremental annual revenue.
And then comes the most important point: The infrastructure investment doesn’t stop after one year. Another generation gets built. Older equipment gets replaced. Power capacity expands. Networking gets upgraded. The capital base compounds. So does the return requirement.
This Is Why Revenue Growth Alone Isn’t Enough
The question investors should ask is not simply whether AI revenue is growing. Of course it is. The harder question is: Is incremental AI-related operating profit growing faster than the capital base required to produce it?
Those are very different tests.
- Imagine AI usage doubling while inference prices fall 60%.
- Token consumption looks extraordinary.
- Infrastructure utilization rises.
- AI becomes more important.
- But revenue only increases modestly.
- Meanwhile, depreciation, electricity, networking and replacement costs continue rising.
Technologically, AI is winning. Economically, returns could still deteriorate. That is how a technology boom becomes a capital bubble.
The Hyperscalers Have One Enormous Advantage
There is a strong counterargument to everything I have written so far. The companies financing much of this buildout are not speculative dot-com startups. They are some of the most profitable corporations ever created. Amazon, Microsoft, Alphabet and Meta possess enormous cash-generating businesses.
- They own distribution.
- They control customer relationships.
- They operate global infrastructure.
- They can spread AI investments across multiple businesses.
That diversification matters enormously.
- Microsoft can monetize AI infrastructure through Azure, Microsoft 365, GitHub, security, Dynamics and other services.
- Google can monetize it through Google Cloud, Search, advertising, Workspace, YouTube and Gemini.
- Amazon can monetize it through AWS while applying AI throughout retail, logistics, advertising and its other businesses.
- Meta doesn’t even need to sell AI compute externally to generate a return. If better recommendation models increase engagement or better advertising models improve conversion, AI infrastructure can monetize through the advertising machine.
This is why comparing hyperscaler capex only against cloud revenue is economically wrong. The infrastructure supports far more than cloud. It also explains why these companies can tolerate overcapacity much longer than smaller AI infrastructure providers. They have balance sheets capable of waiting for demand. That doesn’t eliminate the requirement for returns. It extends the runway.
Circular Capital Is Where Things Get More Complicated
Another characteristic of the AI cycle deserves considerably more scrutiny.
- Capital increasingly moves in circles.
- Hyperscalers invest in model companies.
- Model companies make enormous infrastructure commitments to hyperscalers.
- Semiconductor companies invest across the AI ecosystem.
- Infrastructure companies raise capital partly against contracted AI demand.
- Model builders use that infrastructure to build products that drive additional infrastructure consumption.
None of this is inherently problematic. Much of it makes perfect strategic sense.
- Cloud providers need anchor AI customers.
- Model builders need compute.
- Chip companies benefit from larger ecosystems.
- Infrastructure providers need financing.
But economically, we need to distinguish between ecosystem demand and independent end-market demand. Ultimately, someone outside that capital loop has to pay.
- A bank.
- A manufacturer.
- A healthcare company.
- A government.
- An advertiser.
- A developer.
- A consumer.
Someone has to generate enough incremental economic value from AI to fund the return expectations throughout the infrastructure stack. That is where the bubble question ultimately gets resolved.
The Market-Resolving Event
I don’t believe the AI bubble bursts because AI stops working. The more interesting market-resolving event occurs when supply finally catches demand. Imagine several things happening together.
- GPU availability improves.
- HBM capacity expands.
- Advanced packaging constraints ease.
- New data centers receive power.
- Alternative accelerators mature.
- Custom silicon takes more workloads.
- Inference software becomes dramatically more efficient.
- Model prices continue falling.
- Enterprises become much better at model routing.
On the surface, almost everything on that list is bullish for AI. And technologically, it is. But economically, it removes scarcity. And scarcity is currently preventing us from seeing the industry’s true clearing price. When compute becomes broadly available, we finally discover what customers are actually willing to pay for it. At that point, the industry moves from a capacity race to a utilization and return race. That transition could be brutal for poorly positioned infrastructure providers.
What Would Actually Pop the Bubble?
I don’t expect one dramatic event. I would watch for a combination of six conditions.
1. Supply catches demand. GPU, memory and data center capacity become sufficiently available that customers stop reserving almost everything they can obtain.
2. Inference prices fall faster than workload consumption increases. Token volumes continue exploding while revenue growth begins slowing.
3. Productive utilization disappoints. The industry discovers that installed compute significantly overstates economically monetizable compute.
4. Depreciation catches up with capex. The enormous 2025-2027 infrastructure build begins flowing through income statements at scale.
5. Hardware becomes economically obsolete faster than expected. Older GPUs still function but cannot compete economically against newer systems.
6. Free cash flow deteriorates without a corresponding acceleration in AI monetization.
Meta’s latest quarter provides a fascinating example of what to watch.
- It generated $31.86 billion of operating cash flow.
- It spent $31.08 billion on capital expenditures.
- Free cash flow was just $784 million.
- At the same time, revenue grew 28%.
That is not evidence that Meta has an AI problem. It demonstrates how capital-intensive the AI race has become. If AI investment produces sustained improvements in advertising, engagement and new revenue streams, that spending may prove extraordinarily valuable. If incremental returns eventually fall below the cost of capital, the interpretation changes very quickly.
We’ve Seen This Movie Before
There is a tendency to argue that because AI will transform the economy, today’s investment cannot constitute a bubble.
History says otherwise. The Internet transformed the economy. It also produced an enormous capital bubble. Fiber networks were real. Data centers were real. Internet traffic was real. The productivity improvements were real. And enormous amounts of investor capital were still destroyed.
The mistake wasn’t believing in the Internet. The mistake was assuming technological importance automatically justified every dollar of capital deployed at every valuation and every point in the infrastructure cycle. Much of the infrastructure built during the telecom and Internet bubble eventually became extremely valuable. But the companies and investors that financed it did not necessarily capture that value. That distinction should matter enormously to AI investors. A technology can transform the global economy while simultaneously destroying capital. Those outcomes are not contradictory.
The Four Numbers I Would Watch
Forget the argument over whether AI is “a bubble.” Watch four numbers.
1. Capital deployed
How much incremental infrastructure is actually being built?
2. Productive utilization
How much of that infrastructure is performing economically useful work?
3. Incremental cash flow
How much additional cash generation can reasonably be attributed to AI-enabled products and services?
4. Return on invested capital
After depreciation, replacement costs, power, networking, financing and technological obsolescence, what economic return is the infrastructure actually generating?
Those four numbers will tell us far more about the sustainability of the AI boom than another benchmark leaderboard.
AI Can Change Everything and Still Be Overbuilt
I remain highly bullish on AI as a technology. I am bullish on inference. I am bullish on AI infrastructure. I believe AI will reshape cloud computing, software development, cybersecurity, data infrastructure and the way enterprises operate.
But being bullish on the technology does not require being blind to the economics. That is what makes this cycle so fascinating. Both sides can be right.
- AI demand can explode.
- AWS, Azure and Google Cloud can continue growing.
- NVIDIA can sell extraordinary amounts of compute.
- Memory demand can remain exceptionally strong.
- Agents can create entirely new categories of workloads.
- Enterprises can generate enormous productivity improvements.
- And the industry can still build too much infrastructure, too quickly, at a cost that ultimately produces inadequate returns.
That is not an AI failure. That is capital-cycle economics.
The AI bubble won’t burst because AI stops working. It will burst if the economics stop working. And if you want to know whether that is beginning to happen, don’t start with the benchmark leaderboard. Start with the cash flow statement.

