Evaluating whether large language models (LLMs) reason about morality in human-like ways requires more than measuring whether they produce the right outputs in isolated cases. Existing approaches – including scalar agreement, distributional analysis, rationale classification, and consistency testing – compare model and human responses one case at a time and cannot detect how a model structures its moral judgments. This paper applies and extends Peterson and Gärdenfors's (2024) geometric moral sp…
Read moreEvaluating whether large language models (LLMs) reason about morality in human-like ways requires more than measuring whether they produce the right outputs in isolated cases. Existing approaches – including scalar agreement, distributional analysis, rationale classification, and consistency testing – compare model and human responses one case at a time and cannot detect how a model structures its moral judgments. This paper applies and extends Peterson and Gärdenfors's (2024) geometric moral space framework to evaluate structural value alignment in LLMs, that is, whether the relational structure reconstructed from a model's responses matches the structure reconstructed from human respondents on the same task. Sixteen LLMs across three scale tiers (small, medium, and large) were evaluated on ten moral cases under both zero-shot and chain-of-thought (CoT) prompting. The results reveal phenomena invisible to case-by-case metrics. First, all five large models achieve full geometric alignment on these cases (moral space overlap, MSO = 1.00) in at least one prompting condition, and three do so in both; two medium models also reach MSO = 1.00. Second, structural alignment is bimodal in the sense that models either replicate the human moral space or fall substantially short, with no intermediate cases. Third, smaller models systematically lack a distinct fairness region, a failure of moral categorization rather than of labeling. Fourth, CoT prompting simultaneously improves case-by-case similarity correlations and worsens geometric alignment in smaller models, showing that the two can move in opposite directions. Because structural alignment is measured from model responses, it does not by itself establish that a model holds the values in question. It is, however, a more demanding test than case-by-case agreement, and a better motivated and more informative basis for value alignment measurement.