Abstract
Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models
has seen a commensurate rise in ways these models cause harm, spanning areas from toxic speech and hallucinations to AI agents
executing unauthorized actions. Given that these models are probabilistic and general-purpose by nature, it is impossible to enumerate
all possible uses and outputs of the model to reach a fully aligned end state. Within the field of AI safety, these harmful instances are
often framed as “the alignment problem,” of models being “misaligned” with human values. Researchers have responded by pursuing
applied and theoretical AI “value alignment” efforts, often without specifying what they mean by human values. How does the field of
AI value alignment conceive of human values? How are these conceptions of values technically operationalized and evaluated? What
does the emergent theory of value from this field signify for the future of AI?
The study of human values has long been part of many academic disciplines outside of computer science, yet these disciplines
are seldom consulted in AI alignment. Building on the theoretical insights of Zhi-Xuan’s (2024) "preferentist paradigm" critique, we
conduct a review of influential AI alignment literature. We also draw from conceptions of human values from philosophy, anthropology,
and sociology, to create an analytical schema. We annotated 94 AI value alignment research papers to discern their implicit theory of
values in AI. The majority do not define values, relying heavily on “preferences” as a stand-in that runs the risk of reducing complex,
culturally situated concepts down to binary choices. As researchers dispense with using human annotators for model training and
evaluation, turning instead to synthetic data and LLM-as-a-judge approaches to aligning and evaluating models, we identify the
potential to close off alternative methods for contesting and enacting values in foundation models. Overall, value alignment is often
reduced to an exercise in utility maximization, which we argue abstracts human values away from their lived context. In making AI
value alignment’s philosophical commitments explicit, we seek to bring greater specificity and under-explored perspectives into the
debate on whether and how AI can address human values