Abstract:AI alignment is essential for the safe deployment of advanced AI systems. Given that values and preferences change over time, culture, social roles, and context, we need to develop a better understanding of the possible long-term consequences of AI alignment, in particular considering the likely ubiquitous future use of personalized AI assistants. We introduce a flexible and extensible mathematical modelling framework, rooted in social physics, aimed at answering macro-level questions regarding the evolving social norms in human populations under the assumption of frequent AI use. Our analysis is part-analytical, and part-simulation, enabling us to characterize the long-term dynamical consequences under a diverse set of starting assumptions. We highlight the risk of value lock-in, and normative mode collapse, prominently featured in non-adaptive alignment formulations. Beyond alignment, we advocate for the wider adoption of these kinds of social physics models as an epistemic bridge: enabling rapid, rigorous, and quantitatively-grounded hypothesis testing for sociotechnical foresight in general AI futures, and acting as a tractable precursor to more computationally expensive large-scale agentic evaluations.
Submission history
From: Nenad Tomasev [view email]
[v1]
Mon, 20 Jul 2026 21:02:39 UTC (1,469 KB)