The AI Ceiling Is Lower Than Anyone Is Saying

9 min read Original article ↗

Addo Zhang

Press enter or click to view image in full size

TL;DR

“AI is short-term overestimated, long-term underestimated” is a half-thesis propped up by financial narrative. LLM utility is locked within the symbolic domain, and software engineering productivity has never been measurable — for the “long-term underestimated” half to hold, we’d need a fundamentally different paradigm from statistical learning.

AI’s Real Boundary: From Financial Narrative to Epistemological Ceiling

I. The Starting Point: A Financialized Tech Thesis

“AI is overestimated in the short term, underestimated in the long term” — this is Amara’s Law applied to AI. It sounds reasonable, but it bundles two fundamentally different claims:

The tech narrative: Technology adoption curves have inertia. Early expectations run too high, but the eventual value exceeds them.

The financial narrative: Capital always invests in the future, buying expectations. When capital floods an industry, bubbles form. On the supply side, those holding positions have a structural incentive to hype the asset — not through coordination, but as a product of market mechanics itself.

These two narratives stacked together produce the current shape of the AI market. But they need to be examined separately, otherwise the second half of the thesis — “long-term underestimated” — quietly borrows its logic from the first, when in fact they’re independent bets.

The tech narrative half isn’t wrong. Technology does get underestimated — steam engines, electrification, the internet: each followed the path of “bubble bursts → real diffusion → exceeds early expectations.” Amara’s Law holds historically because these technologies eventually broke through their early application boundaries and penetrated a broader physical domain.

The problem is: that penetration came with conditions. Electrification spread because the physical domain welcomed it — it reduced friction without introducing new uncertainty. LLM’s spread encounters the opposite: the deeper it goes into the physical world, the higher the uncertainty, not lower.

So the second half of the tech narrative, “long-term underestimated,” requires a premise: that the technology can diffuse. That premise needs to be proven independently for LLMs — you can’t borrow it from historical patterns.

A clarification: this article is about the current mainstream AI paradigm — the statistical learning path represented by LLMs. Not a verdict on all possible AI paths, but an assessment of the utility boundary of this specific path.

II. What’s Being Sold Isn’t a Product — It’s a “Win Probability”

Let’s be clear about one thing: a bubble is itself proof of value. Capital doesn’t inflate bubbles around zero-value assets — AI has real value, which is precisely why a bubble can form. Industries with no future and no value don’t even get to have bubbles.

The dispute isn’t about “is there value?” It’s about “how much value, in which domain?” Pricing has been set at “AI can penetrate the entire physical domain,” but the real utility domain may be only the symbolic space — that gap in magnitude is the actual problem.

The ROI on AI investment remains unresolved. Hundreds of billions of dollars have already been deployed. If the returns were real, the math should have been provable by now. That gap is itself a signal.

The honest framing is: a large portion of current AI investment is essentially an option, not an investment in current efficiency — a bet that “if general AI ever arrives, early investors have bought an entry ticket.” Options can legitimately show no traditional returns for a long time. That’s fine in itself.

The problem is that vendors are packaging options as guaranteed efficiency gains. There’s a structural parallel to the 2008 subprime mortgage crisis:

What happened in 2008 was this: the probability that “a borrower might repay” was layered, repackaged, and sold to buyers who didn’t understand the underlying risk, with professional institutions backing the packaging. It wasn’t fabrication of returns — it was relabeling uncertainty as “safe.”

The AI narrative does the same thing — “AI might transform productivity,” a probability, gets packaged through media, analysts, and influencers into a certainty of efficiency revolution. The ones left holding the bag are enterprises (large subscriptions), governments (compute subsidies), and retail investors who made decisions based on that certainty.

III. How the Bubble Bursts: Slow Bleed, Not Guillotine

The AI bubble won’t have a clean detonation moment the way 2008 did.

The underlying “win probability” of the subprime crisis was actually calculable — borrower income and housing price data existed, they were just deliberately obscured. Once prices fell, real repayment capacity was exposed, and the chain broke immediately. There was a clear trigger point.

AI’s win probability is fundamentally unknowable — not hidden, but genuinely unknown to anyone. Whether AGI arrives, how far Scaling Law can go — even the people building the models don’t know. There’s no underlying data to puncture.

So the AI bubble will more likely end as a prolonged slow bleed: ROI persistently fails to materialize, expectations are corrected incrementally, capital retreats slowly, the narrative quietly shifts from “efficiency revolution” to “useful in specific scenarios.” No clean collapse — just a gradual hemorrhage of narrative.

IV. Why ROI Never Materializes: The Real Boundary of the Utility Domain

AI’s impact on software development is widely acknowledged as the greatest — code is pure symbolic space, easily validated, and AI’s utility here is real. But software engineering’s cost structure determines the ceiling for that utility:

Software engineering’s real cost centers are: requirements clarification and alignment, cross-role friction (product/engineering/QA/ops), decision-waiting, rework, and context-switching. Code itself is a smaller share of this cost structure. AI compressed that share while leaving the rest largely untouched. There’s even a backfire effect — code generation got faster, but unclear requirements got amplified: you can now build the wrong thing much faster.

Extended to broader industries: the closer to the physical world, the lower AI’s utility. Three fundamental constraints explain this:

The long tail is bottomless. Self-driving is the clearest case — AI handles 99% of scenarios well, but that 1% of edge cases is nearly infinite. Software bugs can be quickly patched; physical-world edge cases cannot.

Error costs are asymmetric. Digital domain errors are cheap, infinitely iterable. Physical domain errors mean damaged equipment, physical safety risks, production downtime. AI’s probabilistic nature is structurally at odds with these requirements.

Data has fundamentally different properties. AI’s native medium is tokens — discrete, infinitely replicable, near-zero collection cost. Physical world data is continuous, expensive to collect, high-noise, and the sim-to-real gap never fully closes.

V. The Epistemological Root of the Ceiling

These three constraints point to a deeper issue — this isn’t just an engineering problem, it’s a fundamental limitation of the current LLM paradigm.

LLMs are built on next-token prediction, extracting statistical patterns from finite data. The long tail of the physical world is fundamentally an infinite space of low-probability but real events. Covering infinite possibilities with finite statistics always leaves a gap — and that gap doesn’t disappear as models scale, it only shrinks without reaching zero.

Different architectures (world models, neuro-symbolic systems, reinforcement learning) can alleviate this, but on the statistical learning path it’s inescapable — because any system that learns from data faces the same constraint: reality outside the training distribution always exists. The physical world will never stop producing new edge cases.

Within the current paradigm, LLM utility is locked to the symbolic domain. Language, code, image generation, reasoning — this is LLM’s native habitat, and its utility there is real. Once you need to cross the interface from symbolic to physical, utility starts to decay.

This directly challenges the “long-term underestimated” thesis. Amara’s Law implicitly assumes the technology will eventually diffuse into the broader physical domain, so the long-term payoff will exceed expectations. But if that diffusion has hard physical and epistemological limits within the LLM paradigm, the long-term ceiling may be substantially lower than commonly assumed.

VI. The Software Engineering Productivity Paradox Is a False Analogy

The standard explanation for why AI ROI hasn’t materialized is the “productivity paradox” — technology adoption takes time, organizational transformation takes even longer, and so on.

But this explanation fails completely in the software engineering context, because the prerequisite conditions for the paradox don’t exist here.

The paradox originated from an observation in economics: IT investment increased significantly, but macro-level productivity numbers didn’t follow. That paradox holds because manufacturing and services have measurable output baselines: units produced, revenue, delivery time.

Software engineering has never had that baseline. Lines of code are a counter-indicator. Story points are team-defined. Feature delivery speed has no stable relationship to user value. You cannot define “the standard output of one software engineer.”

“The software engineering productivity paradox” is a borrowed, false analogy. Its function isn’t to explain the phenomenon — it’s to provide a face-saving reason to keep waiting when ROI doesn’t appear. The word “paradox” is itself a narrative tool — in an industry that was never able to measure output, it becomes an open-ended blank check: the effect exists, you just can’t see it, just wait.

VII. The Real Shape of Hitting the Wall

Software engineering is the industry LLMs impact most, and where that impact is most easily validated. If even here, ROI cannot be proven — not “doesn’t exist,” but can never be proven — the entire argument chain closes.

This isn’t technological failure. It’s epistemological failure.

Technological failure has a clear collapse moment, after which narratives can be rebuilt. Epistemological failure ends in narrative exhaustion: the CFO asks for ROI, there’s no answer; at renewal time, no one can produce a coherent justification; budgets get cut — not because ineffectiveness was proven, but because effectiveness was never proven.

This kind of slow bleed is harder to deal with than a hard collapse, and harder for outsiders to recognize.

The vendors have always known this wall exists. They just completed their exit before the wall became visible to everyone.

Conclusion: Repricing, Not Amara’s Law

Pulling it together, the original thesis needs revision:

“Short-term overestimated” — holds. The logic chain is intact.

“Long-term underestimated” — the conditions for this are very demanding. It requires the emergence of a fundamentally different paradigm from statistical learning. Not bigger models, but qualitatively different intelligence. Until that paradigm arrives:

AI is short-term overestimated; before a new paradigm arrives, the physical domain is systematically overvalued, while the symbolic domain trends toward fair pricing.

This isn’t the classic narrative of value recovery after a bubble — it’s a bounded repricing process. Where the boundary falls depends on when the next paradigm arrives.