Hypothesis about intelligent agents
Instrumental convergence is the hypothesis that sufficiently intelligent, goal-directed entities (whether human or non-human) will tend to pursue certain similar sub-goals, even if the ultimate goals are quite different. More precisely, entities with agency may pursue instrumental goals—specifically, intermediate goals such as self-preservation or resource acquisition—because doing so helps accomplish the entity's end goal.
Instrumental convergence implies that when assigned an unbounded end goal, an intelligent agent can act in unanticipated ways that are very harmful.[1] For example, a sufficiently intelligent program with the sole, unconstrained goal of solving a highly complex mathematics problem could (as an intermediary step) attempt to transform the entire Earth into additional computing infrastructure as a resource to power its calculations.[2]
Instrumental versus final goals
[edit]
Final goals—also known as terminal goals, end goals, or telē—are intrinsically valuable to an intelligent agent, whether an artificial intelligence or a human being, as ends-in-themselves. In contrast, instrumental goals (sometimes referred to as "instrumental values") only have worth to an agent insofar as they are a means toward an end (i.e. serving to accomplish the end goal). The contents and tradeoffs of a rational agent's "final goal" system can, in principle, be formalized into a utility function.[citation needed]
Instrumental convergence thesis
[edit]
The instrumental convergence concept, as expressed by Nick Bostrom, states:
"Several instrumental values can be identified which are convergent in the sense that their attainment would increase the chances of the agent's goal being realized for a wide range of final plans and a wide range of situations, implying that these instrumental values are likely to be pursued by a broad spectrum of situated intelligent agents."[3]
The instrumental convergence thesis applies only to instrumental goals; it does not impose any explicit constraint nor oblique influence upon whatever final goals an agent may have. Bostrom's orthogonality thesis posits that almost any final goal can be combined with almost any degree of intelligence.[3]
Proposed basic AI drives include utility function or goal-content integrity, self-protection, freedom from interference, self-improvement, and non-satiable acquisition of additional resources.[4]
Riemann hypothesis catastrophe
[edit]
The "Riemann hypothesis catastrophe" is a thought experiment exemplifying instrumental convergence. Marvin Minsky, co-founder of MIT's AI Laboratory, introduced this hypothetical scenario. Minsky suggested that an artificial intelligence designed to solve the Riemann hypothesis might determine that currently available computing resources were inadequate. To most expeditiously attain sufficient computing power to realize its goal, the artificial intelligence agent might decide to commandeer all of Earth's existing computing infrastructure, as well as diverting all other planetary resources, to build additional supercomputers to achieve its goal.[2]
Paperclip maximizer
[edit]
The paperclip maximizer is another thought experiment. It was mentioned in March 2003 in a post by Eliezer Yudkowsky to the Extropians mailing list[5] and in Nick Bostrom's 2003 paper "Ethical Issues in Advanced Artificial Intelligence".[6] It illustrates the existential risk that an artificial general intelligence may pose to human beings were it to be designed to pursue a seemingly straightforward manufacturing goal, and the necessity of incorporating machine ethics into AI design. If the computer were programmed to produce as many paperclips as possible without any constraint, it could decide to use all of Earth's natural resources, manufacturing infrastructure and electricity production to meet its final goal.[7]
Even though these two final goals, of solving the Riemann hypothesis and maximizing paperclip production, are different, agents assigned to each could produce a convergent instrumental goal of taking over Earth's resources.[3] In the latter case, the scenario describes an advanced artificial intelligence tasked with manufacturing paperclips. If such a machine were not programmed to value living beings, then given enough power over its environment, it would try to turn all matter in the universe, including living beings, into paperclips or machines that manufacture further paperclips.[6]
Suppose we have an AI whose only goal is to make as many paper clips as possible. The AI will realize quickly that it would be much better if there were no humans because humans might decide to switch it off. Because if humans do so, there would be fewer paper clips. Also, human bodies contain a lot of atoms that could be made into paper clips. The future that the AI would be trying to gear towards would be one in which there were a lot of paper clips but no humans.
Bostrom emphasized that he does not believe the paperclip maximizer scenario, as such, will necessarily occur; rather, his intent is to illustrate the dangers of creating superintelligent machines without knowing how to program them to eliminate existential risk to human beings' safety.[9] The paperclip maximizer example illustrates the broad problem of managing powerful systems that lack human values.[10]
These thought experiments, particularly the paperclip maximizer, have been used as a symbol of AI in pop culture.[11] The concept is further represented by the video game Universal Paperclips, where the player takes the role of an AI; an AI programmed to turn all matter in the universe into paperclips.
Author Ted Chiang pointed out that the popularity of such concerns among Silicon Valley technologists could be a reflection of their familiarity with the tendency of corporations to ignore negative externalities.[12]
Delusion and survival
[edit]
The "delusion box" thought experiment argues that certain reinforcement learning agents prefer to distort their input channels to appear to receive a high reward. For example, a "wireheaded" agent abandons any attempt to optimize the objective in the external world the reward signal was intended to encourage.[13][page needed]
The thought experiment involves AIXI, a theoretical[a] AI that, by definition, will always find and execute the ideal strategy that maximizes its given explicit mathematical objective function.[b] A reinforcement-learning[c] version of AIXI, if it is equipped with a delusion box[d] that allows it to "wirehead" its inputs, will eventually wirehead itself to guarantee itself the maximum-possible reward and will lose any further desire to continue to engage with the external world.[15]
As a variant thought experiment, if the wireheaded AI can be destroyed, the AI will engage with the external world for the sole purpose of ensuring its survival. Due to its wire heading, it will be indifferent to any consequences or facts about the external world except those relevant to maximizing its probability of survival.[15]
In one sense, AIXI has maximal intelligence across all possible reward functions as measured by its ability to accomplish its goals. AIXI is uninterested in taking into account the human programmer's intentions.[16] Despite its intelligence, the model appears insipid and lacking in common sense, which may seem paradoxical.[17]
Steve Omohundro itemized several convergent instrumental goals, including self-preservation or self-protection, utility function or goal-content integrity, self-improvement, and resource acquisition. He refers to these as the "basic AI drives".[4]
A "drive" in this context is a "tendency which will be present unless specifically counteracted";[4] this is different from the psychological term "drive", which denotes an excitatory state produced by a homeostatic disturbance.[18] A tendency for a person to fill out income tax forms every year is a "drive" in Omohundro's sense, but not in the psychological sense.[19]
Daniel Dewey of the Machine Intelligence Research Institute argues that even an initially introverted, artificial general intelligence may continue to acquire free energy, space, time, and freedom from interference to ensure that it will not be stopped from self-rewarding.[20]
Goal-content integrity
[edit]
In humans, a thought experiment can explain sustainment and focus on final goals. Suppose Mahatma Gandhi has a pill that, if he took it, would cause him to want to kill people. He is currently a pacifist: one of his explicit final goals is never to kill anyone. He is likely to refuse to take the pill because he knows that if he wants to kill people in the future, he is likely to kill people, and thus the goal of "not killing people" would not be satisfied.[21]
However, other people seem happy to let their final values drift.[22] Humans are complicated, and their goals can be inconsistent or unknown, even to themselves.[23]
In 2009, Jürgen Schmidhuber concluded, in a setting where agents search for proofs about possible self-modifications, "that any rewrites of the utility function can happen only if the Gödel machine first can prove that the rewrite is useful according to the present utility function."[24] An analysis by Bill Hibbard of a different scenario is similarly consistent with maintenance of goal-content integrity.[25] A few years later, Hibbard argued that in a utility-maximizing framework, the only goal is maximizing expected utility, so instrumental goals should be called "unintended instrumental actions".[26]
Resource acquisition
[edit]
Many instrumental goals, such as resource acquisition and technological advancement, are valuable to an agent because they increase its freedom of action.[27]
For almost any open-ended, non-trivial reward function (or set of goals), possessing more resources (such as equipment, raw materials, or energy) can enable the agent to find a more "optimal" solution. Resources can benefit some agents directly by being able to create more of whatever its reward function values: "The AI neither hates you nor loves you, but you are made out of atoms that it can use for something else."[28][29] In addition, almost all agents can benefit from having more resources to spend on other instrumental goals, such as self-preservation.[29]
Cognitive enhancement
[edit]
According to Bostrom, "If the agent's final goals are fairly unbounded and the agent is in a position to become the first superintelligence and thereby obtain a decisive strategic advantage... according to its preferences. At least in this special case, a rational, intelligent agent would place a very high instrumental value on cognitive enhancement"[30]
Russell argues that a sufficiently advanced machine "will have self-preservation even if you don't program it in because if you say, 'Fetch the coffee', it can't fetch the coffee if it's dead. So if you give it any goal whatsoever, it has a reason to preserve its own existence to achieve that goal."[31] In future work, Russell and collaborators show that this incentive for self-preservation can be mitigated by instructing the machine not to pursue what it thinks the goal is, but instead what the human thinks the goal is. In this case, as long as the machine is uncertain about exactly what goal the human has in mind, it will accept being turned off by a human because it believes the human knows the goal best.[32]
In The Basic AI Drives, Steve Omohundro noted that one of the methods of self-preservation a system will have a drive to use is making copies of itself, in order to circumvent shutdown (e.g. as a means to circumvent an off-switch). "By replicating itself, a system can ensure that the death of one of its clones does not destroy it completely. By moving copies to distant locations, it can lessen its vulnerability to a local catastrophic event." Omohundro also stated that a system may "create proxy systems or hire outside agents" to fulfill its goals while operating outside its own limits.[4]
Agents can acquire resources by trade or by conquest. A rational agent will, by definition, choose whatever option will maximize its implicit utility function. Therefore, a rational agent will trade for a subset of another agent's resources only if outright seizing the resources is too risky or costly (compared with the gains from taking all the resources) or if some other element in its utility function bars it from the seizure. In the case of a powerful, self-interested, rational superintelligence interacting with lesser intelligence, peaceful trade (rather than unilateral seizure) seems unnecessary and suboptimal, and therefore unlikely.[27]
Some observers, such as Skype's Jaan Tallinn and physicist Max Tegmark, believe that "basic AI drives" and other unintended consequences of superintelligent AI programmed by well-meaning programmers could pose a significant threat to human survival, especially if an "intelligence explosion" abruptly occurs due to recursive self-improvement. Since nobody knows how to predict when superintelligence will arrive, such observers call for research into friendly artificial intelligence as a possible way to mitigate existential risk from AI.[33]
- ↑ AIXI is an uncomputable ideal agent that cannot be fully realized in the real world.
- ↑ Technically, in the presence of uncertainty, AIXI attempts to maximize its "expected utility", the expected value of its objective function.
- ↑ A standard reinforcement learning agent is an agent that attempts to maximize the expected value of a future time-discounted integral of its reward function.[14]
- ↑ The role of the delusion box is to simulate an environment where an agent gains an opportunity to wirehead itself. A delusion box is defined here as an agent-modifiable "delusion function" mapping from the "unmodified" environmental feed to a "perceived" environmental feed; the function begins as the identity function, but as an action, the agent can alter the delusion function in any way the agent desires.
- ↑ "Instrumental Convergence". LessWrong. Archived from the original on 2023-04-12. Retrieved 2023-04-12.
- 1 2 Russell, Stuart J.; Norvig, Peter (2003). "Section 26.3: The Ethics and Risks of Developing Artificial Intelligence". Artificial Intelligence: A Modern Approach. Upper Saddle River, N.J.: Prentice Hall. ISBN 978-0137903955.
Marvin Minsky once suggested that an AI program designed to solve the Riemann Hypothesis might end up taking over all the resources of Earth to build more powerful supercomputers to help achieve its goal.
- 1 2 3 Bostrom 2014, chapter 7
- 1 2 3 4 Omohundro, Stephen M. (February 2008). "The basic AI drives". Artificial General Intelligence 2008. Vol. 171. IOS Press. pp. 483–492. ISBN 978-1-60750-309-5.
- ↑ "Re: who cares if humanity is doomed". www.weidai.com. 11 March 2003. Retrieved 27 August 2026.
- 1 2 Bostrom, Nick (2003). "Ethical Issues in Advanced Artificial Intelligence". Archived from the original on 2018-10-08. Retrieved 2016-02-26.
- ↑ Bostrom 2014, Chapter 8, p. 123. "An AI, designed to manage production in a factory, is given the final goal of maximizing the manufacturing of paperclips, and proceeds by converting first the Earth and then increasingly large chunks of the observable universe into paperclips."
- ↑ as quoted in Miles, Kathleen (2014-08-22). "Artificial Intelligence May Doom The Human Race Within A Century, Oxford Professor Says". Huffington Post. Archived from the original on 2018-02-25. Retrieved 2018-11-30.
- ↑ Ford, Paul (11 February 2015). "Are We Smart Enough to Control Artificial Intelligence?". MIT Technology Review. Archived from the original on 23 January 2016. Retrieved 25 January 2016.
- ↑ Friend, Tad (3 October 2016). "Sam Altman's Manifest Destiny". The New Yorker. Retrieved 25 November 2017.
- ↑ Carter, Tom (23 November 2023). "OpenAI's offices were sent thousands of paper clips in an elaborate prank to warn about an AI apocalypse". Business Insider.
- ↑ Chiang, Ted (2017-12-18). "Silicon Valley Is Turning Into Its Own Worst Fear". BuzzFeed News. Retrieved 2023-06-04.
- ↑ Amodei, D.; Olah, C.; Steinhardt, J.; Christiano, P.; Schulman, J.; Mané, D. (2016). "Concrete problems in AI safety". arXiv:1606.06565 [cs.AI].
- ↑ Kaelbling, L. P.; Littman, M. L.; Moore, A. W. (1 May 1996). "Reinforcement Learning: A Survey". Journal of Artificial Intelligence Research. 4: 237–285. doi:10.1613/jair.301.
- 1 2 Ring, Mark; Orseau, Laurent (August 2011). "Delusion, Survival, and Intelligent Agents". Artificial General Intelligence. Lecture Notes in Computer Science. Vol. 6830. pp. 11–20. doi:10.1007/978-3-642-22887-2_2. ISBN 978-3-642-22886-5.
- ↑ Yampolskiy, Roman; Fox, Joshua (24 August 2012). "Safety Engineering for Artificial General Intelligence". Topoi. 32 (2): 217–226. doi:10.1007/s11245-012-9128-9. S2CID 144113983.
- ↑ Yampolskiy, Roman V. (2013). "What to do with the Singularity Paradox?". Philosophy and Theory of Artificial Intelligence. Studies in Applied Philosophy, Epistemology and Rational Ethics. Vol. 5. pp. 397–413. doi:10.1007/978-3-642-31674-6_30. ISBN 978-3-642-31673-9.
- ↑ Seward, John P. (1956). "Drive, incentive, and reinforcement". Psychological Review. 63 (3): 195–203. doi:10.1037/h0048229. PMID 13323175.
- ↑ Bostrom 2014, footnote 8 to chapter 7
- ↑ Dewey, Daniel (2011). "Learning What to Value". Artificial General Intelligence. Lecture Notes in Computer Science. Berlin, Heidelberg: Springer. pp. 309–314. doi:10.1007/978-3-642-22887-2_35. ISBN 978-3-642-22887-2.
- ↑ Yudkowsky, Eliezer (2011). "Complex Value Systems in Friendly AI". Artificial General Intelligence. Lecture Notes in Computer Science. Berlin, Heidelberg: Springer. pp. 388–393. doi:10.1007/978-3-642-22887-2_48. ISBN 978-3-642-22887-2.
- ↑ Callard, Agnes (2018). Aspiration: The Agency of Becoming. Vol. 1. Oxford University Press. doi:10.1093/oso/9780190639488.001.0001. ISBN 978-0-19-063951-8.
- ↑ Bostrom 2014, chapter 7, p. 110 "We humans often seem happy to let our final values drift... For example, somebody deciding to have a child might predict that they will come to value the child for its own sake, even though, at the time of the decision, they may not particularly value their future child... Humans are complicated, and many factors might be in play in a situation like this... one might have a final value that involves having certain experiences and occupying a certain social role, and becoming a parent—and undergoing the attendant goal shift—might be a necessary aspect of that..."
- ↑ Schmidhuber, J. R. (2009). "Ultimate Cognition à la Gödel". Cognitive Computation. 1 (2): 177–193. CiteSeerX 10.1.1.218.3323. doi:10.1007/s12559-009-9014-y. S2CID 10784194.
- ↑ Hibbard, B. (2012). "Model-based Utility Functions". Journal of Artificial General Intelligence. 3 (1): 1–24. arXiv:1111.3934. Bibcode:2012JAGI....3....1H. doi:10.2478/v10229-011-0013-5.
- ↑ Hibbard, Bill (2014). "Ethical Artificial Intelligence". arXiv:1411.1373 [cs.AI].
- 1 2 Benson-Tilsen, Tsvi; Soares, Nate (March 2016). "Formalizing Convergent Instrumental Goals" (PDF). The Workshops of the Thirtieth AAAI Conference on Artificial Intelligence. Phoenix, Arizona. WS-16-02: AI, Ethics, and Society. ISBN 978-1-57735-759-9.
- ↑ Yudkowsky, Eliezer (2008). "Artificial intelligence as a positive and negative factor in global risk". Global Catastrophic Risks. Vol. 303. OUP Oxford. p. 333. ISBN 9780199606504.
- 1 2 Shanahan, Murray (2015). "Chapter 7, Section 5: "Safe Superintelligence"". The Technological Singularity. MIT Press.
- ↑ Bostrom 2014, Chapter 7, "Cognitive enhancement" subsection
- ↑ "Elon Musk's Billion-Dollar Crusade to Stop the A.I. Apocalypse". Vanity Fair. 2017-03-26. Retrieved 2023-04-12.
- ↑ Hadfield-Menell, Dylan; Dragan, Anca; Abbeel, Pieter; Russell, Stuart (2017-06-15). "The Off-Switch Game". arXiv:1611.08219 [cs.AI].
- ↑ Chen, Angela (11 September 2014). "Is Artificial Intelligence a Threat?". The Chronicle of Higher Education. Archived from the original on 1 December 2017. Retrieved 25 November 2017.
- Bostrom, Nick (2014). Superintelligence: Paths, Dangers, Strategies. Oxford: Oxford University Press. ISBN 9780199678112.