I haven’t been shy about voicing my negative opinions of the current large language model hype. But I’m hesitant to describe myself as “anti-AI”. And it’s not because I have some sort of super-nuanced position like “LLMs are fine in some cases” or “I’d be OK with LLMs if only we could solve these problems”. I’m in fact adamantly opposed to the very assumptions on which the hype is based: that a statistical model of word co-occurrence frequency is a good foundation for a multi-purpose tool for answering questions or generating meaningful text or code. Train a local model on text acquired with the creators’ express permission and run it on a machine fueled entirely by renewable energy, and I still don’t think it’s a good idea. The fact that commercial LLMs disregard consent and copyright and consume ungodly amounts of resources only makes a fundamentally bad thing even worse.
So why don’t I call myself “anti-AI”? Well, because of two ambiguities: one in the definition of “AI”, and another in what it means to be “anti-AI”. And that’s why I’m proposing an alternate nomenclature to explain exactly what my position is.
Why “AI” is a mostly meaningless term
The term “artificial intelligence” is usually believed to have come into existence around the time of the Dartmouth Summer Research Project on Artificial Intelligence in 1956. Even then, it wasn’t exactly clear what the term referred to; it was chosen as a sort of blanket term for a lot of semi-related research in information theory, computer science, and other disciplines. Broadly speaking, it seemed to refer to attempts to mimic the effects of human thought with computers. I say “mimic the effects of human thought” rather than “mimic human thought,” because the people involved were under no illusion that the processes by which computers handled knowledge representation, language processing, visual recognition, or other “AI” endeavors was anything like what was happening in the human brain. Absolutely nobody thought that human brains had Lisp interpreters running in them to parse all the sentences we hear. That’s why they used the word “artificial”: to stress that this was an imitation of a thing, not the thing itself. (The belief that brains really do some form of computation similar to what goes on inside a digital computer, and that we can design computer systems that actually mimic the human brain, belongs to the related but separate discipline of cognitive science.)
In the beginning, most things that were called “artificial intelligence” were rule-based systems, not essentially different from other types of computer programs. The main difference was that while computing had traditionally been used for tasks that were relatively simple to express in algorithmic form, such as calculating projectile trajectories or solving algebraic equations, artificial intelligence sought to conquer less obviously algorithmic territory, such as representing the information content of natural-language text, reasoning about it, and forming a natural-language reply. But over time, “machine learning” became a thing: using automated statistical or quasi-statistical methods to generalize from a set of input data. This might involve sorting pictures of food into the broad categories “hot dog” or “not hot dog,” or predicting what a website customer might want to buy next given the last several things they’ve bought. This represented a very different approach than what came to be known as “good old-fashioned artificial intelligence” (GOFAI): instead of working out the steps needed to perform a task and telling the computer to do them, you just threw a bunch of data at a program and let the computer figure out what to do. At least, that’s how it’s often anthropomorphically described. In fact, it’s just the same rule-based approach at a higher level of abstraction: instead of rules for turning inputs into outputs, it’s higher-order rules for turning training data into rules for turning inputs into outputs. But it’s a lot easier to say the computer “learns” how to do things, which unfortunately convinces a lot of people that machine learning systems are doing something a lot more similar to the human brain than they really are.
So now “artificial intelligence” was being used to refer to both old-school algorithms for manipulating symbols, and new-fangled statistical trickery. Two very different technologies described with the same terminology, their only commonality being a vague desire to emulate what people do. Over time, as computing hardware got cheaper and more powerful, “deep learning” became a thing. This was really just a multi-layer perceptron, a slightly more advanced version of a machine learning method dating back to the 1950s. Only now, we could make them much bigger, and train them on much bigger datasets. At around the same time, the Internet was really taking off, making it easier than ever to scoop up terrifying amounts of data, especially natural-language data. Even though deep learning wasn’t a fundamentally different kind of thing than earlier forms of machine learning, the scale involved allowed us to do a lot of things (image recognition, automatic translation) that earlier methods were only sort of OK at. So a lot of people started viewing deep learning as its own thing, making yet one more category under the ever-expanding umbrella of “AI”.
Most recently, people realized that if you take a certain kind of deep learning architecture, and an enormous amount of text, you can train it to predict what the next word will be given several hundred preceding words. And they wrapped a framework around this that would simulate dialogue, with a human providing one half of the conversation and the word-predictor providing the other. And it was so good at producing human-sounding dialogue, and the venture capital money was so plentiful, that companies started shoving it into every software product on the market. But they didn’t call it “large language models” or “next-token predictors” or even “chatbots”. They just called it “AI”.
The result is that now, when somebody says “AI,” there’s about a 99% chance they’re referring to one specific application of one particular form of one of several broad categories of software that have all, at various points in time, been called “AI”. And this one particular sub-sub-sub-discipline of AI is quite unpopular, for its environmental impact, its unethically sourced training data, the fact that employers keep bragging that it will let them replace human workers, and the fact that it isn’t actually very good at most of the things it’s sold to do. So lots and lots of people are saying “I hate AI!”
The inherent ambiguity of the term then allows the same tech CEOs who have been so quick to associate the term “AI” with their narrow selection of products the perfect comeback: “You hate AI? But look at all these wonderful things AI has done! Look at how your phone can recognize people in your images and tag them! Look at how it can replace your picture with a cartoon cat in real-time! Look at how doctors can identify cancer cells that would previously have gone undetected! If you don’t like AI, you must hate all these things too!” And, of course, none of these things use LLMs in any way. They were all around before tech companies started shoving chatbots into everything. They are not, in any way, shape, or form, what the vast majority of “AI haters” are talking about when they say “AI,” and the tech bros know it. They just don’t care, because twisting language to make it more difficult to express dissent has long been a favorite trick of authoritarians. Double plus ungood.
Well, I’m not falling for it. Wherever possible, I’m trying to avoid using the term “artificial intelligence” or “AI” at all, and instead say exactly what I’m talking about (LLMs, logistic regression, support vector machines, expert systems, Lisp, whatever.)
“Anti-AI” isn’t very meaningful either
So the ambiguity of the term “AI” lets LLM-shillers attack their opponents by accusing them of opposing things they’re not actually opposed to. But what if we simply redefined “AI” to mean “large language models”? Could we then call ourselves “anti-AI”?
Well, no, because LLM boosters have astroturfed that term. There are organizations like PauseAI out there that might seem to be anti-AI, but have the support of some of the biggest AI company CEOs. And when you look closer at their complaints, there’s scarcely any mention of environmental concerns, intellectual property, de-skilling, or other concerns that people actually have about LLMs. Instead, the argument is that “chatbots are super-duper smart and getting exponentially smarter! It’s only a matter of time before they try to take over the world and turn us into paperclips! So we have to pause all AI development and set up strict regulations and global governing bodies to make sure AI development, which is inevitable and can never be stopped, is done the right way!”
This argument has two goals. One, it takes it as given that LLMs are extremely intelligent and getting ever smarter, with the hopes that people will just accept it, and move past the question of “are these things even smart?” straight to “what do we do about these incredibly smart machines?” The premise, however, is demonstrably false. While the precise definition of intelligence is tough to pin down, some things just fail to meet even the broadest criteria for intelligence. Large language models do not learn. They undergo a one-time “training phase,” during which the weights between their perceptron nodes are modified based on training data. Once that phase is over, and before the network is ever deployed to production, the weights are fossilized in amber and will never change again. The only way to get more “information” into a multi-layer perceptron is to train a new network with more training data to replace it. Sure, chatbots might seem to remember things you tell them, but that’s just a parlor trick: every time you query a chatbot, the entire history of the conversation so far (or at least the entire history that will fit inside its context window) is fed back to the network. An LLM is really just a function that takes a very huge binary number (the encoded conversation) and returns a much smaller binary number (corresponding to the predicted next token). It has no inner life. It doesn’t think about anything in between prompts, or in fact change at all. To call anything it does “intelligence” insults the very concept of intelligence. So of course, asking “is this thing intelligent?” is something the LLM-shillers never, ever want you to do.
The other goal of the “Pause AI” argument is more prosaic: regulatory capture. If today’s big AI companies can convince the world’s governments to enact AI regulations that the AI companies help write, and form a government body that AI companies get to help pick, they can ensure that they get to do whatever they want to do, while simultaneously raising entry costs for potential competitors.
Actually, there might be a third goal, or at least a happy unintended consequence: the idea that chatbots can become super-powerful and take over the world is so patently ridiculous that most people would dismiss it out of hand, and rightfully so. If LLM-shillers can take control over the narrative so that “anti-AI” comes to be synonymous with “believes in Skynet but with an odd paperclip fixation,” they can make any opposition sound ridiculous.
I must admit, I had my QUALLMS
In order to distinguish “opposition to LLMs” from “opposition to random-forest classifiers,” and to distance the legitimate concerns about LLMs from the astroturfed anti-Skynet hype, I’ve come up with an alternative to “anti-AI”: I call it QUALLMS: Quit Using All Large Language Models. It has the following advantages:
- It specifically targets LLMs, rather than vague “artificial intelligence”
- It makes clear what we want to happen (quit using them, rather than enact Anthropic-friendly regulation)
- It’s catchy
(The association with Widow’s Bay is just an added bonus.)