•  2
    Evaluating whether large language models (LLMs) reason about morality in human-like ways requires more than measuring whether they produce the right outputs in isolated cases. Existing approaches – including scalar agreement, distributional analysis, rationale classification, and consistency testing – compare model and human responses one case at a time and cannot detect how a model structures its moral judgments. This paper applies and extends Peterson and Gärdenfors's (2024) geometric moral sp…Read more