•  45
    Large language models (LLMs) are increasingly attributed with internal representations of the world, a claim often grounded in the teleosemantic account of representation. I examine the core properties this account requires and argue that pre-trained and fine-tuned LLMs face significant challenges in satisfying the selection-history criterion, whereas agentic LLMs may not. I proceed in three stages. First, I argue that in purely pre-trained LLMs, activation vectors and features represent the sta…Read more
  •  203
    I contend that, when making claims about LLMs’ world models, we should ask whether a model represents not only the surface properties of an input but also its underlying causes and principles. I then support, drawing on empirical studies in mechanistic interpretability, that features and circuits are strong candidates for realizing both kinds of representation. However, I argue that the pragmatics of human linguistic production can hinder both forms of representation in text-trained LLMs, althou…Read more