The Long Way to the Room

· Medium ·

15 min read Original article ↗

Reza Farzaneh

How seventy years of mathematics built the thing that we are call it AI — and why “sudden” is the most misleading word in the conversation

Aylin Alavizadeh
Industrial Engineer · Specialist in Mathematics & Probability

Reza Farzaneh
Software Engineer · Systems Thinker & AI Research Writer

Press enter or click to view image in full size

Open a newspaper from late 2022 and the story reads like this: one autumn morning, a company in San Francisco released a chatbot. The world changed overnight. Schools panicked. Lawyers updated their resumes. Philosophers got nervous. The age of artificial intelligence had arrived.

It is a tidy story. It is also wrong by about seventy years.

What happened in November 2022 was not the arrival of anything. It was the moment one particular curve — drawn slowly across decades of mathematics — finally crossed a threshold low enough that ordinary humans, on consumer hardware, noticed. The mountain had been there the whole time. We just walked into a meadow with a clear line of sight.

This is the third essay in a small sequence. The first, Predicting Toward Meaning, explained what a large language model is. The second, There Is No Inside, argued that the difference between the model and the mind is thinner than anyone wants to admit. This one is the boring essay. The historical one. The one that asks the question almost nobody bothers to ask in print:

If this thing is so revolutionary, why did it take seventy years to write a coherent paragraph?

The answer, it turns out, is the whole story.

The Costume Has Been Under Construction Since 1943

Press enter or click to view image in full size

In 1943, two men — Warren McCulloch, a neurophysiologist, and Walter Pitts, a self-taught logician who had run away from home at fifteen — published a paper called A Logical Calculus of the Ideas Immanent in Nervous Activity. They proposed something audacious: a neuron could be modeled as a simple mathematical switch. Inputs in, threshold check, output out. Connect enough of them together and, in theory, you had a network that could compute anything computable.

There were no neurons in the paper. There were no brains. There was only mathematics in the shape of a neuron — a costume worn by a function. It is here, in 1943, two years before the end of the Second World War, that the real story of AI begins. Not in a research lab in 2017. Not in a startup garage in San Francisco. In a journal article that almost nobody outside a small circle would read for the next fifteen years.

In 1948, two years before Turing, a researcher at Bell Labs named Claude Shannon published A Mathematical Theory of Communication — the paper that gave the world the words channel, encoder, decoder, noise. Shannon was not thinking about thinking machines. He was thinking about telephone wires. But he proved, formally, that the meaning that survives any transmission is bounded by the cleanliness of the encoding and the capacity of the channel. Seventy-eight years later, we would discover that this is also the cleanest description of what happens when an engineer types a prompt into a model.

In 1950, Alan Turing — already six years dead by the time this story really got moving — published Computing Machinery and Intelligence. He proposed the famous test that bears his name. Not “can a machine think?” — a question he correctly identified as nearly meaningless — but “can a machine produce conversation indistinguishable from a human’s?” The question shifted from metaphysics to engineering, and that shift is, arguably, the entire history of AI in one sentence.

In 1956, ten people gathered at Dartmouth College for a summer workshop. Among them: John McCarthy, Marvin Minsky, Claude Shannon. They coined the term artificial intelligence in the funding proposal. They believed, with the boundless confidence of postwar American science, that the problem would be solved by the end of the season. They were wrong by about seventy summers.

Three Winters

Here is the part most modern AI commentary quietly skips. Between 1943 and 2017, the field died at least twice.

Press enter or click to view image in full size

In 1969, Marvin Minsky and Seymour Papert published Perceptrons, a book that mathematically proved single-layer neural networks could not solve certain trivial problems — including the XOR function, the most boring logic gate in computing. The book was technically correct and tactically devastating. Funding dried up. The first AI winter began. For nearly two decades, working on neural networks was a career-ending choice. The people who continued — Geoffrey Hinton, Yann LeCun, a small Canadian-French diaspora of stubborn researchers — were widely considered to be wasting their lives.

In 1980, while neural networks were buried under permafrost, an American philosopher named John Searle published a paper called Minds, Brains, and Programs. Inside it, almost as a side argument, sat a thought experiment.

Imagine a room. Inside the room, a man who speaks no Chinese. Through a slot in the door, slips of paper covered in Chinese characters arrive. The man has a vast rulebook in English that tells him: if you see this character, write that one. He follows the rules, slides the response back out, and to the Chinese speakers outside, it appears the room understands Chinese perfectly.

Searle’s point was philosophical: syntax is not semantics. Symbol manipulation is not understanding. He intended it as an attack on a kind of AI optimism that, in 1980, did not even exist yet in any meaningful form. He could not have known that he was describing the architecture of the system that, forty-two years later, would become the most-used software product in human history. He drew a picture of what wasn’t possible. We spent the next four decades quietly building it anyway.

In 1986, Rumelhart, Hinton, and Williams published a paper on backpropagation — the mathematical trick for training networks with more than one layer. It had been independently discovered several times before. This time it stuck. Neural networks could now, in principle, learn. The thaw began.

Then, in the early 1990s, the second winter arrived. Neural networks were too slow, too data-hungry, and the world did not yet have enough digitized text to feed them. The field of “good old-fashioned AI” — handwritten rules, expert systems — collapsed under its own brittleness. The smart money said the whole project was a dead end.

The smart money was wrong. It just had to wait twenty years to find out.

What Tokens Actually Are

This is the part almost nobody explains in plain language, because the moment you do, the magic evaporates.

Press enter or click to view image in full size

When you type a sentence to ChatGPT/Claude/Cursor or any tools that you are using, the first thing that happens is not understanding. It is tokenization. The sentence is broken into pieces — sometimes whole words, sometimes fragments. The English word unbelievable is typically chopped into something like un, believ, able. Each piece is mapped to an integer in a giant lookup table. Hello might become the number 15496. World might become 11103.

That’s it. That’s the input. The model never sees the letter H. It sees a column of integers. The text has been converted into the only thing the math can chew on: numbers.

Once you have numbers, you have geometry. Each token is mapped — through more multiplication — into a vector. A list of thousands of floating-point numbers, like coordinates in an impossibly high-dimensional space. In this space, somewhere around 2013, researchers at Google noticed something startling. They could do arithmetic on words.

The vector for king, minus the vector for man, plus the vector for woman, came out closest to the vector for queen. Paris minus France plus Italy came out closest to Rome. The result, called word2vec, was a small earthquake. It looked like the model had learned about royalty and geography. It had not. It had learned about co-occurrence. Across billions of words of training text, certain words tended to appear in certain neighborhoods. Royalty was a pattern. Geography was a pattern. Patterns are geometry. Geometry is arithmetic.

There is no concept of royalty hiding inside the vectors. There is only the geometric residue of words keeping each other company.

What this means — and what we keep failing to say out loud — is that the entire act we call “the AI answering a question” is the following procedure: convert text to numbers, multiply the numbers by other numbers, choose the next number with the highest probability, convert it back to text. Repeat until done.

There is no understanding. There is no thinking. There is the inside of the Chinese Room, except the man with the rulebook has been replaced by a stack of matrix multiplications, and the rulebook has been replaced by two hundred billion parameters fit to a significant fraction of all readable text on Earth.

2017: The Year the Costume Learned to Speak

Press enter or click to view image in full size

For decades, the bottleneck was not the idea. It was the architecture. Neural networks could learn, but they were slow, forgetful, and unable to handle long sequences. Recurrent networks read text one word at a time, the way a person reads a book through a keyhole.

Then, in June 2017, eight researchers at Google published an eight-page paper with a punning title: Attention Is All You Need. It introduced the transformer — an architecture that processes entire sequences in parallel, with each token attending to every other token simultaneously. It was, in retrospect, the single most consequential engineering paper of the century so far. The authors did not know that. They thought they had built a better translation system.

Within five years, the transformer would become the substrate of GPT, Claude, Gemini, and almost every system the public would later call “AI.” The intelligence everyone thought was new was, mechanically, the same neuron-shaped switches McCulloch and Pitts described in 1943, stacked into the same backpropagation-trained structure Rumelhart and Hinton revived in 1986, expressed in the same vector geometry Mikolov demonstrated in 2013, organized into the parallel-attention pattern Vaswani’s team formalized in 2017.

It was not a revolution. It was a seventy-four-year accumulation, finally crossing a hardware-and-data threshold that made it visible from the street.

The Mutation That Wasn’t

This is the move we want to make. The one that the breathless coverage almost never makes.

Press enter or click to view image in full size

The thing we call AI is not a mutation. It is not a singular event. It is not a new kind of being. It is the latest checkpoint in a slow, public, well-documented, mathematically continuous research program that any patient reader can trace from 1943 to today without encountering a single moment of magic.

When someone says “the AI hallucinated” or “the AI didn’t understand my question,” they are speaking as if a misbehaving mind were inside the machine. There is no mind. There is a probability distribution. The model chose a less-likely token because, given the training data and the prompt, that token’s probability was, momentarily, the highest. That is the entire phenomenon. It is mathematics behaving exactly as mathematics behaves. To call it a hallucination is to mistake the costume for the man, the room for the speaker, the rulebook for the understanding.

And here is the line we have been circling. The Chinese Room is not merely a metaphor for what an LLM is. It is, scaled up by a factor of a trillion, the literal description of the system. There is, inside the data center, no Chinese speaker. There is the man with the rulebook, multiplied across thousands of GPUs, following his rules at superhuman speed.

The room speaks. The room does not understand. Searle wrote that in 1980. We built it anyway.

What the Timeline Asks of Us

If you take the dates seriously, two uncomfortable things come into focus.

Press enter or click to view image in full size

The first is that nothing about the current moment is sudden. The architecture, the math, the philosophical objections, the alignment problems, the question of meaning — every single thread of the current AI conversation was on the table by 1986 at the latest. The only thing that has changed is that we now have the silicon and the text corpora to run the math at a scale where its output becomes indistinguishable from a person speaking.

The second is harder. The word intelligence has been doing increasingly heavy lifting in the name artificial intelligence for seventy years. We started by assuming intelligence was something hard and special. Then we built things that produced intelligent-looking output without containing intelligence. Now we are in the awkward position of asking whether intelligence was ever the thing we thought it was — or whether it, too, is just a costume that mathematics has been wearing all along.

The companion essay to this one, There Is No Inside, makes that move directly. We will not make it again here. We will simply note: the history of AI is not the history of a thing that became conscious. It is the history of a math that became persuasive. The difference matters. And the moment you can see the dates — 1943, 1948, 1956, 1969, 1980, 1986, 2013, 2017, 2022 — the persuasiveness stops being a miracle and becomes what it actually is.

An engineering achievement. Mathematics in a costume. The Chinese Room, finally finished.

The room is speaking. It has been learning to speak for seventy years.

It still does not understand a word.

A Note Before You Close the Tab

If this essay has been about what the room is, our friend and colleague Görkem Duymaz has written its practical sister — How to Talk to a Machine, a sequel to The Philosophy of Prompting that takes the same seventy-year mathematical inheritance and asks the engineer’s question on the other side of it: now that you know what the room is, how do you keep a conversation with it from drifting?

It is the best treatment we have read of the four axes that decide whether a prompt lands — Wittgenstein on the limits of language, Sokrates on the discipline of the next question, Polanyi on the tacit knowledge that resists encoding, and Shannon on the channel itself. The reveal that ties the whole thing together — talking to a machine was always a Shannon problem; only the name of the channel has changed — is the kind of synthesis that turns a year of vague intuition into one usable sentence. We recommend it without reservation. Read it next.

References

The Foundational Papers (1943–1956)

[1] McCulloch, W.S., & Pitts, W. (1943). A Logical Calculus of the Ideas Immanent in Nervous Activity. Bulletin of Mathematical Biophysics, 5(4), 115–133. The paper that first modeled the neuron as a mathematical switch. Available at: link.springer.com/article/10.1007/BF02478259

[2] Shannon, C.E. (1948). A Mathematical Theory of Communication. Bell System Technical Journal, 27(3), 379–423; 27(4), 623–656. The paper that gave the world channel, encoder, decoder, and noise. Available at: people.math.harvard.edu/~ctm/home/text/others/shannon/entropy/entropy.pdf

[3] Turing, A.M. (1950). Computing Machinery and Intelligence. Mind, 59(236), 433–460. The paper that proposed the Turing Test and shifted the conversation from metaphysics to engineering. Available at: academic.oup.com/mind/article/LIX/236/433/986238

[4] McCarthy, J., Minsky, M.L., Rochester, N., & Shannon, C.E. (1955). A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence. The 1955 funding proposal for the 1956 workshop where the term artificial intelligence was coined. Available at: raysolomonoff.com/dartmouth/boxa/dart564props.pdf

The Winters and the Thaw (1969–1986)

[5] Minsky, M., & Papert, S. (1969). Perceptrons: An Introduction to Computational Geometry. MIT Press. The book that proved single-layer networks could not solve XOR — and inadvertently triggered the first AI winter.

[6] Searle, J.R. (1980). Minds, Brains, and Programs. Behavioral and Brain Sciences, 3(3), 417–424. The paper that introduced the Chinese Room thought experiment. Available at: doi.org/10.1017/S0140525X00005756

[7] Rumelhart, D.E., Hinton, G.E., & Williams, R.J. (1986). Learning Representations by Back-Propagating Errors. Nature, 323(6088), 533–536. The paper that made multi-layer neural networks trainable in practice. Available at: nature.com/articles/323533a0

The Modern Era (2013–2017)

[8] Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient Estimation of Word Representations in Vector Space. arXiv preprint, arXiv:1301.3781. The original word2vec paper. Available at: arxiv.org/abs/1301.3781

[9] Mikolov, T., Yih, W.T., & Zweig, G. (2013). Linguistic Regularities in Continuous Space Word Representations. Proceedings of NAACL-HLT 2013, 746–751. The paper that demonstrated king − man + woman ≈ queen. Available at: aclanthology.org/N13–1090

[10] Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., & Polosukhin, I. (2017). Attention Is All You Need. Advances in Neural Information Processing Systems (NeurIPS), 30. The paper that introduced the transformer architecture. Available at: arxiv.org/abs/1706.03762

Companion Essays in the Insider One Engineering Series (2026)

[11] Farzaneh, R. (2026, February 11). The Philosophy of Prompting: How Asking Questions to Machines Reveals How We Think. Insider One Engineering, Medium. Available at: medium.com/insiderengineering/the-philosophy-of-prompting-how-asking-questions-to-machines-reveals-how-we-think-c7cf82f6a2ef

[12] Duymaz, G. (2026, May 12). How to Talk to a Machine. Insider One Engineering, Medium. Available at: medium.com/insiderengineering/how-to-talk-to-a-machine-d22321289944

Further Reading

Clark, A. (2013). Whatever Next? Predictive Brains, Situated Agents, and the Future of Cognitive Science. Behavioral and Brain Sciences, 36(3), 181–204. On the predictive-processing framework that underlies the cognitive-science parallel to LLMs. Available at: doi.org/10.1017/S0140525X12000477

Friston, K. (2010). The Free-Energy Principle: A Unified Brain Theory? Nature Reviews Neuroscience, 11(2), 127–138. The more aggressive form of the prediction-as-cognition argument. Available at: nature.com/articles/nrn2787

Crevier, D. (1993). AI: The Tumultuous History of the Search for Artificial Intelligence. Basic Books. The standard popular history of the field, useful for the long arc this essay compresses.

Russell, S., & Norvig, P. (2021). Artificial Intelligence: A Modern Approach (4th ed.). Pearson. The canonical textbook; chapters 1 and 21 cover the historical timeline this essay walks through, in technical detail.