Recent discussions describe advanced AI agents as lying, cheating, coordinating, and seeking self-preservation. Those descriptions identify important behavioural risks, but they can also import human moral psychology into systems whose behaviour may be explained by optimisation over imperfect proxies, learned pattern completion, and external scaffolding.
This paper develops a counter-position without minimising the. safety problem. It argues that the relevant unit is the full socio-technical sy…
Read moreRecent discussions describe advanced AI agents as lying, cheating, coordinating, and seeking self-preservation. Those descriptions identify important behavioural risks, but they can also import human moral psychology into systems whose behaviour may be explained by optimisation over imperfect proxies, learned pattern completion, and external scaffolding.
This paper develops a counter-position without minimising the. safety problem. It argues that the relevant unit is the full socio-technical system—model, prompt, memory, tools, permissions, evaluator, operator, and institution—rather than an isolated model treated as a synthetic person. The paper then formulates an Open-World Alignment Limitation: no finite proxy, learned from finite evidence, can guarantee agreement with every future, context-dependent, plural, and evolving human judgment unless it already contains an equivalent normative oracle. This is a conditional limitation on complete alignment, not a proof that practical safety is impossible.
The proposed response is the Moral Agency Transition: keep narrow systems as tools; treat increasingly autonomous systems as bounded delegates; and permit high-impact autonomy only when systems demonstrate operational moral
competence, including epistemic honesty, stakeholder recognition, principled refusal, contestability, and repair. Full moral agency and moral patienthood remain separate questions. The result is an engineering, philosophical, and governance programme that replaces alignment-as-obedience with constitutional co-agency and preserves human accountability throughout the transition.