In 1965, Gordon Moore had a handful of data points and a hunch: the number of components on an integrated circuit would keep growing exponentially. The cadence later became known as a doubling roughly every two years, and it held for decades. But I doubt anybody in that era really understood what they were looking at. They thought they were arguing about how many components you could etch onto silicon. Few could have imagined that, 60 years later, transistors would become the substrate of the world — every phone, car, thermostat, pacemaker, and pair of earbuds. The transistor didn't just get smaller. It got everywhere.
The scale is hard to hold in our heads. A leading chip went from ~2,300 transistors in 1971 to 58 billion in 2021. We now manufacture close to 10²¹ of them a year — more than the industry made in every year before 2017 combined. The transistor is the most-produced object in human history.
I think we're watching the same movie again — except this cut runs at roughly 8x speed.
The new unit isn't the transistor. It's the token: the unit by which we measure AI consumption, the thing a model ingests and emits. And the curve looks eerily familiar. Comparable data is hard to find, but Google has disclosed some figures in its I/O presentations. In early 2024, Google was processing 9.7 trillion tokens a month across its products. A year later, 480 trillion. A year after that, 3.2 quadrillion. That's about 330x in two years — and that's one company.
Transistors doubled every couple of years. Tokens are doubling every few months. What took transistors seventeen years of compounding, tokens are doing in two.
The analogy differs in two important ways.
First, speed: the transistor explosion unfolded across the mainframe, the PC, and then the smartphone. The token explosion is compressing a comparable magnitude of adaptation into just a few years.
Second, scope: a transistor is a fixed physical component, while a token is a unit of computation. Both face physical and economic constraints, so the analogy is imperfect. But tokens can be spent against an open-ended set of language, reasoning, and software tasks.
If token costs keep falling and availability keeps rising, products will stop treating inference as a scarce, user-triggered event. They will spend it continuously — adapting interfaces, checking work, and handling small tasks that are uneconomical today.
So here's the bet I'd make: tokens will end up where transistors did — invisible, ambient, and embedded throughout products rather than confined to a few obviously “AI” features. The teams positioned for that future will build as if tokens are already infinite: not because physical constraints disappear, but because products designed around today's scarcity will age badly as costs fall.
The transistor took fifty years to vanish into the substrate. Tokens may get there within the next decade.