LLM capabilities are increasing. They cause risks to financial and logistical infrastructure, and make engineered pandemics more likely. So we should work on safety of these models and also societal mitigation efforts. Maybe even pause (?).
But in this post I will address an emotional part of human disempowerment that comes with smarter AI systems.
Computers can do things better than we can, and this hurts our ego. It sucks that we can no longer be uniquely good at something. When humans were the most intelligent things around, if you were unskilled, you could at least dream that one day you could great. But now, even the most skilled of us are unlikely to win against the machines on any timescale. Scientists and mathematicians are demoralised. So are artists.
This is hurtful, partly because we consider computers separate from us. The computer is a different thing, separate from our sense of self. This separate entity removes a source of happiness, like the feeling of being good at something. This makes us feel bad and disempowered. Another word for this is competition. We are in competition with the machines, and we have lost. The language we use to describe this is “Y has beaten me at X”.
But could I change the framing so that it feels better? Could I reframe so that I am excited about the prospects of LLM intelligence? If so, I might feel more empowered.
One language game might be to consider LLMs as part of our being. Right now, we have an idealogical boundary between LLMs and ourselves. Could the feeling of disempowerment be solved by removing this idealogical line?
Why do English football fans, many of which are bitter rivals during the English Premier League, suddenly all band together during the World Cup? They consider themselves as one entity, and remove the idealogical lines from local teams during World Cup years.
Most people America felt good when the first moon landing happened. Amongst Americans, there was no competition nor adversity, despite Republican and Democrat divides. Why? American humans considered themselves as one entity. For a moment, America was Neil Armstrong (except the Soviets, who considered themselves as separate).
Suppose that I removed the idealogical boundary between me and and an LLM. In this paradigm, the LLM becomes as an extension of me. I merge with the LLM. If I do this, will I feel less bad about my inadequacy against the machine? How can I compete with something that I consider part of myself?
Maybe this is the solution to the disempowerment. If I prompt an LLM to do something great, then the emotional win is for both the LLM and I. I cease to feel bad.
There is one complication though. The ideological merge almost surely involves a third party, an unwanted dinner guest. This third party is the model provider; OpenAI, Anthropic, Deepmind, or whoever. With a third party, there are many bad things that can be injected in the merge. They might inject their values in unsavoury ways. They could invade your privacy. The LLM might run away and imitate you in ways that you don’t want.
Therefore, the only acceptable merge would be merge with a local LLM on a local GPU cluster. This local LLM would reflect my values and beliefs and enhance them in ways that I feel comfortable with. Right now, there is no such service for me to design my own local aligned LLM. I don’t have enough compute. I have been experimenting with models that can run on Mac M4 Studio memory but the results have been bad.
I need a way to fully specify my values and beliefs, and transfer them fully to the local LLM. One way could be having my LLM learn continuously from my Obsidian and blog, and then have it ask me about my beliefs to fill in the gaps. For unspecified moral problems, maybe it could extrapolate from what I have already.
I don’t know how much data it would need though.
This is because mathematically, my ideals and beliefs from a very high dimensional space. The high dimensional space changes, and its unclear what the principal eigenvectors of this high dimensional space is. The high dimensional space gets projected down to my writing, and this writing is supposed to reflect my values.
For an LLM to be fully aligned, it would need to take the low dimensional space of my writing and then recreate the high dimensional space that is my complete value set. But this is a problem of information loss. The idea of recreating a high dimensional space from low dimensional space has deep connections with complex systems and theoretical physics. The holographic principle in string theory relates information in lower dimensional boundaries to a higher dimensional bulk. In complex systems research people try to recreate dynamics from a reduced representation called phase space.
But its unclear how successful this approach would be in an LLM and values context.
For example, would my feelings on capitalism be enough for an LLM to infer my feelings on abortion? Would my feelings on veganism be enough for an LLM to infer whether I want to buy that new toaster? Is this problem just a facade of the alignment problem that everyone’s working on right now? It’s all unclear.
Acknowledgements
Ideas sparked from an Independent Science Society dinner with Ariel Cheng. All views and mistakes my own.