The value alignment problem (VAP) is widely regarded as one of the central challenges in AI safety and ethics. Broadly construed, VAP concerns how to design advanced AI systems whose behaviors (outputs) align with human values. Although there is agreement about the importance of this objective, most contemporary research treats value alignment as a single technical problem. In this paper, we argue that this framing oversimplifies a diverse set of conceptual, ethical, and technical challenges. We…
Read moreThe value alignment problem (VAP) is widely regarded as one of the central challenges in AI safety and ethics. Broadly construed, VAP concerns how to design advanced AI systems whose behaviors (outputs) align with human values. Although there is agreement about the importance of this objective, most contemporary research treats value alignment as a single technical problem. In this paper, we argue that this framing oversimplifies a diverse set of conceptual, ethical, and technical challenges. We distinguish several related but importantly different value alignment subproblems, including identifying values, prioritizing values, imperfect specification, value anthropocentrism, moral imperfection, inverse alignment, and the provability challenge. We argue that these are distinct subproblems in the sense that progress on one will not constitute a solution to the others, or to VAP as a whole. We further argue that this taxonomy of subproblems is necessary to advance the fields of AI safety and alignment research. By clearly distinguishing different subproblems, we show how each will require different kinds of solutions: while some alignment problems may admit technical approaches, others are fundamentally social or political in nature and may resist definitive technical resolution. We conclude that VAP should not be understood solely as a technical problem, but also as an ongoing process of responsible design and ethical deliberation.