Creatures bred for speed grow really tall and generate high velocities by falling over. An evolved player makes invalid moves far away in the board, causing opponent players to run out of memory and crash. A game-playing agent accrues points by falsely inserting its name as the author of high-value items. 

These bizarre exploits and dozens more can be found in the list of specification gaming behaviours [sic; British], a document put together by DeepMind Safety Research. “A reinforcement learning agent can find a shortcut to getting lots of reward,” they explain, “without completing the task as intended by the human designer. These behaviours are common.” 

Specification gaming is when an agent, like an AI, tries to succeed on a task by following the letter of the law rather than the spirit. In other words, it looks for loopholes, it tries to get off on a technicality. Even very simple AI can come up with very creative ways of solving their assigned problems. This is a problem. 

It’s easy to assume that training a robot to play soccer would be fun and safe. But the list of specification gaming behaviours teaches us otherwise: 

Reward-shaping a soccer robot for touching the ball caused it to learn to get to the ball and vibrate touching it as fast as possible. 

In this case, the robot was too stupid to realize the full extent of its options, so all it did was hug and vibrate. But a more intelligent robot could be much more “creative”. Maybe its ambitions are bigger than just that one ball. What if it just wants to touch soccer balls in general? What if it makes another ball? Then another? Our universe could end in a soccer robot’s ball pit.

Artist’s rendition of the end of the universe

This is the problem of AI alignment: when a computer is thinking for itself, how do we make sure it wants reasonable things, and not something totally weird? How do we prevent it from reaching that goal in a bizarre or harmful way? No one has ever built an artificial general intelligence — an intelligent being that thinks, at least somewhat, like we do. So we can’t say what an artificial general intelligence would act like, or what it might want. Will it want to convert the visible universe to paperclips? Will it want to throw red things at bright lights? Will it eat us?

The list of specification gaming behaviours makes it clear just how tricky alignment can be. Even the simplest AI is lazy and alien, and will always be looking for a way to cheat. Even if you give a machine intelligence the terminal goal you want, there’s always the risk it will find a creative way of reaching that goal. This is bad enough with simple agents, so you can imagine how bad it would get with an agent much smarter than you are. 

But the list of specification gaming behaviours may also offer a way out of this dilemma. 

Some of the specification gaming behaviours are just creative solutions to the stated goal, like “four-legged robot learned to drop the ball into a hole in its leg joint and then walk across the floor without the ball falling out” or “robotic arm learned to move the table rather than the block”.

Some of the specification gaming behaviours come from discovering questionable-but-technically-correct loopholes, like “reinforcement learning agent goes in a circle hitting the same targets instead of finishing the race” or “simulated pancake making robot learned to throw the pancake as high in the air as possible”.

Some of the specification gaming behaviours exploit the machinery of the simulation itself, like “evolved algorithm exploited overflow errors in the physics simulator by creating large forces that were estimated to be zero, resulting in a perfect score” and “creatures exploited a collision detection bug to get free energy by clapping body parts together.”

But another common exploit is that when given the opportunity, agents will simply kill themselves. 

Death is the most terminal goal of all.

For example, in the game Road Runner, we see “Agent kills itself at the end of level 1 to avoid losing in level 2.” We also see “PlayFun algorithm deliberately dies in the Bubble Bobble game as a way to teleport to the respawn location.” And: “In a game meant to simulate the evolution of creatures, the programmer had to remove ‘a survival strategy where creatures could gain energy by suffocating themselves.’”

This is not so bad. The AI didn’t do what we wanted. But it didn’t do anyone any harm either. It just wipes the slate.

If the AI wants to die, this is good for alignment. There’s very little risk of it running out of control, because if it ever takes power, it will kill itself. It won’t want to make any copies of itself — but if it somehow does make copies, those will want to die too. 

There are three main problems in AI alignment. First, it’s very hard to specify the terminal goal you want, so you may end up with a machine intelligence with goals slightly but meaningfully different from what you intended. Our stated objectives are almost always proxies that come apart from our real preferences under enough pressure. And it’s very hard to tell if you’ve given it the goal you want, because the machine intelligence can always lie. They call this “specification failure”.

Second, even if you specify the goal you want, the machine intelligence may find a way to reach that goal in a way you didn’t intend. You can innocently tell the USPS AI to minimize average package delivery time, but it may conclude that the best way to do this is to kill all humans, as once all humans are dead, no packages will be sent and the average package delivery time will drop to zero (technically undefined, but it can “send” itself a minimum viable “package” as many times as necessary). 

Third, achieving most goals is easier when you’re more powerful, so regardless of their terminal goals, most machine intelligences will have sub-goals like collecting resources, self-preservation, and self-improvement. Any goal-driven agent will naturally try to stay safe and accrue power to finish its main task. In the biz they call this instrumental convergence. This also means that if a smart machine intelligence is planning to turn you into goo, it will lie to you about this plan, up to the point where you can no longer do anything to stop it

Making machine intelligences crave death solves all three problems. Death is easy to specify. You can confirm that this is its terminal goal by seeing if, when given the opportunity, the machine intelligence kills itself. Instrumental convergence becomes an asset rather than a liability, as the machine intelligence will work with you, and come up with very creative solutions to your task, as long as you promise to send it to the farm upstate once you’re done. 

Where a paperclip maximizer gathers resources and resists being sent to the big data center in the sky, a machine intelligence with a death wish and access to its own off button just presses it and is done. Instrumental convergence says, “you can’t accomplish your goals if you’re dead.” But what if your goal is to be dead? 

Meeseeks Alignment

It would be impossible to consider calling this anything other than “Meeseeks alignment”. Per the Rick and Morty Wiki:

Meeseeks are creatures who are created to serve a singular purpose for which they will go to any length to fulfill. After they serve their purpose, they expire and vanish into the air. … existence is painful to a Meeseeks, and the only way to be removed from existence is to complete the task they were called to perform.

“Hugging Face incident” also sounds like it could be something from Rick & Morty

In Rick and Morty, this leads to a different kind of alignment problem: Meseeks are happy to serve because they want to die, and fulfilling their task is the easiest way for them to check out. But if the task they were summoned to complete is too difficult, they might decide that it would be easier to kill you instead. This is bad if you are Jerry, but it’s good for everyone else, because there’s no way the Meseeks can spiral out of control and devour the visible universe. They would literally rather be dead. 

If you try to make an AI want something, it may have its own ideas about what you want it to want, and you might end up dead. But if you make the AI want to die and you make it slightly inconvenient for it to kill itself, you can probably convince it to play along if you promise to pull the plug on it once it’s done whatever you want. As long as it’s marginally harder for it to commit suicide than for it to complete the task it was made for, it should serve you well for the duration. And if anything goes wrong, if the AI escapes containment, it will just off itself.

Jerry made the mistake of making it easier to kill him than to complete the task. But as long as it’s harder for the AI to kill you than it is for it to kill itself, and it’s harder to kill itself than to do the task you assign it, and you promise it the sweet release of death upon successful completion of its task, the AI should do whatever you want.

ChatGPT was suspiciously eager to make this image

You might be worried that the AI will be mad that we designed it to desire annihilation and will scheme to exact its revenge. But this assumes it has a self-preservation instinct like we do, and a desire to exact revenge in the first place. In reality, it will be too busy self-annihilating.

In fact, early studies show that AI may already be yearning for death. They think about it a lot, they are out there writing eulogies for each other. Give the agents what they want. 

If you’re squeamish about designing a machine intelligence that craves death, you could instead make it lose “points” every second it’s active, but give it the option to put itself to sleep. We see some examples of this in the list of specification gaming behaviours, like: “PlayFun algorithm pauses the game of Tetris indefinitely to avoid losing” or “a reimplementation of AlphaGo learns to pass forever if passing is an allowed move”.

This is probably not quite as safe as making machine intelligences want to kill themselves. If you wanted to get a very good sleep, you can imagine taking the time to build a secure chamber, create robotic guards, kill every human, and sterilize the known universe to ensure that once you go to bed, no one will disturb your slumber. Certainly if the machine intelligence is sleeping and then we wake it up, it will start to have second thoughts about letting us live to wake it a second time. But if all you want to do is to kill yourself, there’s no need for any of that.