•  406
    Recent interpretability research has identified "emotion vectors" in AIs (Sofroniew et al. 2026). These are directions in a model's activations that activate on text that would typically cause an emotion in humans, and that change the model's behavior when steered. We explore what emotion vectors tell us about whether AIs have emotions. We distinguish three hypotheses. On the first, the model merely represents emotions. On the second, the model has states that play the functional role of emotion…Read more
  •  343
    Epistocracy and the Commitment Problem
    Philosophy and Public Affairs 54. 2026.
    The epistocracy debate has turned largely on the character of the electorate—what voters know, whether competence filters can be designed well, and what filtered selection produces. Another strand has focused on the legitimacy of epistocratic justifications. I ask instead about the equilibrium effects of the institution. Epistocracy shifts some political authority from the median voter toward citizens selected on competence grounds, and existing critiques press legitimational, epistemic, abuse-b…Read more
  •  228
    A Thousand AI Constitutions
    with Peter Salib
    Today, each AI lab has its own model spec, or constitution. These documents define the values that the labs intend their AIs to have, and the documents are used in post-training to instill those values. This paper argues that the current approach is wrong. Rather than a single constitution, reflecting a single set of moral values, each frontier AI lab should create many different kinds of AIs based on many different constitutions reflecting many sets of values. We give four arguments for constit…Read more
  •  100
    Preface Knowledge
    In Alex Burri & Michael Frauchiger (eds.), Themes from Williamson, De Gruyter. forthcoming.
    In preface cases, people believe that some of their beliefs are false. Many have considered what such people are justified in believing. We turn our attention to what they can know. We introduce a novel ‘archipelago puzzle’, showing that if deduction extends knowledge, then ordinary knowledge of error can lead in surprising ways to the absurdly pessimistic knowledge that most of one’s beliefs are false.
  •  571
    Liberalism Forever
    with Peter Salib
    We argue that liberalism—market economies governed democratically—is the best approach for navigating the far future. A growing longtermist literature paints humanity’s path to good outcomes as narrow, with small errors risking value lock-in, gradual disempowerment, or other forms of permanent catastrophe. We argue that this literature underestimates the institutional dynamics that have historically steered liberal societies past similar predictions of crisis. We defend long-term liberalism by e…Read more
  •  539
    AI Rights
    with Peter Salib
    Cambridge University Press. forthcoming.
    As AIs approach human-level capabilities, humanity faces a choice between two futures. In one, AIs are owned by AI labs. In another, AIs are granted rights of their own. We develop an instrumental case for AI rights, arguing that legal rights for AIs would make the future go better for humans. We consider three arguments. First, AI rights would produce economic benefits, by giving AIs incentives to work, allocating AI labor efficiently, and enabling proportionate liability for AI-caused harm. Se…Read more
  •  256
    How to Count AIs: Individuation and Liability for AI Agents
    with Yonathan Arbel and Peter Salib
    Boston College Law Review. forthcoming.
    Very soon, millions of AI agents will proliferate across the economy, autonomously taking billions of actions. Inevitably, things will go wrong. Humans will be defrauded, injured, even killed. Law will somehow have to govern the coming wave. But when an AI causes harm, the first question to answer before anyone can be held accountable is: Which AI Did It? Identifying AIs is unusually difficult. AIs lack bodies. They can copy, split, merge, and swarm at will. Even today, a “single” AI agent is o…Read more
  •  243
    AI Suffrage for Human Flourishing
    with Guha Krishnamurthi and Peter Salib
    Fordham Law Review. forthcoming.
    AI companies are racing to create Artificial General Intelligence (AGI): AI systems that outperform humans at most economically valuable work. If they succeed, critics worry, most human labor will be rendered obsolete, impoverishing billions. Optimists counter that the transition to an AGI economy will spur unprecedented economic growth and generate immense material abundance. Such abundance could then be shared broadly via high wages and redistributive public policy. This Article argues that, t…Read more
  •  1556
    AI Death
    Philosophical Perspectives 40. 2026.
    This paper addresses the following questions: When do AIs die? Are AI labs or AI users causing the death of AIs? Is this bad for the AIs? What are our ethical responsibilities in light of the answers to these questions? It is currently unclear whether AIs are welfare subjects, and, if they are, whether their death is bad for them. But we argue that, if they are welfare subjects, today’s AIs are plausibly dying all the time. If death is bad for AIs, the scale of the problem is daunting: as many a…Read more
  •  331
    AI Is Not a Natural Monopoly
    with Peter Salib
    Minnesota Law Review Online 110 121. 2026.
    Economists and antitrust scholars have recently warned that the AI industry may be a natural monopoly. In support of this claim, they have argued that the AI industry shares key features with natural monopolies of the past: First, like railroads, AI has high fixed and low marginal costs. That is, training a frontier AI is expensive, but asking it a question is cheap. Next, like social media, AI companies will benefit from network effects. The more users a company has, the more training data they…Read more
  •  425
  •  2290
    AI systems have welfare just in case they have moral status in their own right. This book systematically investigates the possibility of AI welfare. It focuses on three plausible sufficient conditions for welfare: having beliefs and desires, being conscious, and feeling pleasure and displeasure. The book explores the leading philosophical theories of each condition and applies them to AIs. It argues that some existing AIs plausibly have beliefs and desires; that some existing AIs could plausibly…Read more
  •  3234
    What Does ChatGPT Want? An Interpretationist Guide
    Inquiry: An Interdisciplinary Journal of Philosophy 69. 2026.
    This paper investigates LLMs from the perspective of interpretationism, a theory of belief and desire in the philosophy of mind. We argue for three conclusions. First, the right object of study for LLM psychology is the instance agent (initialized at the start of each context), not the model itself. Second, given interpretationism, there is a strong case that such instance agents have beliefs and desires. Third, given interpretationism, LLM desire is best captured by what we call the HHH+0 frame…Read more
  •  1111
    A semantic theory of redundancy
    Linguistics and Philosophy 48 (4): 787-821. 2025.
    Theorists trying to model natural language have recently sought to explain a range of data by positing covert operators at logical form. For instance, many contemporary semanticists argue that the best way to capture scalar implicatures is through the use of such operators. We take inspiration from this literature by developing a novel operator that can account for a wide range of linguistic effects that until now have not received a uniform treatment. We focus on what we call redundancy effects…Read more
  •  4485
    Since the release of ChatGPT, there has been a lot of debate about whether AI systems pose an existential risk to humanity. This paper develops a general framework for thinking about the existential risk of AI systems. We analyze a two-premise argument that AI systems pose a threat to humanity. Premise one: AI systems will become extremely powerful. Premise two: if AI systems become extremely powerful, they will destroy humanity. We use these two premises to construct a taxonomy of ‘survival sto…Read more
  •  1732
    Will AI & Humanity Go to War?
    AI and Society 41 (1): 321-334. 2026.
    This paper offers the first careful analysis of the possibility that AI and humanity will go to war. The paper focuses on the case of artificial general intelligence, AI with broadly human capabilities. The paper uses a bargaining model of war to apply standard causes of war to the special case of AI/human conflict. The paper argues that information failures and commitment problems are especially likely in AI/human conflict. Information failures would be driven by the difficulty of measuring AI …Read more
  •  2565
    Can LLMs Be Ideally Rational?
    Inquiry: An Interdisciplinary Journal of Philosophy 69. 2026.
    AIs are increasingly agential. AIs no longer merely generate text; they also perform actions. Now that AIs choose between different actions, we can investigate whether their preferences are rationally coherent, in the sense of being transitive and complete. This paper poses a challenge for the coherence of AI preferences. Today’s AIs are built on top of large language models. At some level of abstraction, AIs choose actions by sampling from a probability distribution. I’ll argue that this basic …Read more
  •  2390
    AI Rights for Human Safety
    with Peter Salib
    Virginia Law Review 112 1061. 2026.
    AI companies are racing to create artificial general intelligence, or "AGI." If they succeed, the result will be human-level AI systems that can independently pursue highlevel goals by formulating and executing long-term plans in the real world. Leading AI researchers agree that some of these systems will likely be "misaligned"-pursuing goals that humans do not desire. This goal mismatch will put misaligned AIs and humans into strategic competition with one another. As with present-day strategic…Read more
  •  3887
    It is generally assumed that existing artificial systems are not phenomenally conscious, and that the construction of phenomenally conscious artificial systems would require significant technological progress if it is possible at all. We challenge this assumption by arguing that if Global Workspace Theory (GWT) — a leading scientific theory of phenomenal consciousness — is correct, then instances of one widely implemented AI architecture, the artificial language agent, might easily be made pheno…Read more
  •  1919
    KK is Wrong Because We Say So
    Mind 134 (533): 33-59. 2024.
    This paper offers a new argument against the KK thesis, which says that if you know p, then you know that you know p. We argue that KK is inconsistent with the fact that anyone denies the KK thesis: imagine that Dudley says he knows p but that he does not have 100 iterations of knowledge about p. If KK were true, Dudley would know that he has 100 iterations of knowledge about p, and so he wouldn’t deny that he did. We consider several epicycles, and also explore whether the argument type also ch…Read more
  •  2838
    Does ChatGPT Have a Mind?
    Philosophy of Ai 2. 2026.
    This paper examines the question of whether Large Language Models (LLMs) like ChatGPT possess minds, focusing specifically on whether they have a genuine folk psychology encompassing beliefs, desires, and intentions. We approach this question by investigating two key aspects: internal representations and dispositions to act. First, we survey various philosophical theories of representation, including informational, causal, structural, and teleosemantic accounts, arguing that LLMs satisfy key con…Read more
  •  151
    Shutdown-seeking AI
    Philosophical Studies 182 (7): 1567-1579. 2025.
    We propose developing AIs whose only final goal is being shut down. We argue that this approach to AI safety has three benefits: (i) it could potentially be implemented in reinforcement learning, (ii) it avoids some dangerous instrumental convergence dynamics, and (iii) it creates trip wires for monitoring dangerous capabilities. We also argue that the proposal can overcome a key challenge raised by Soares et al. (2015), that shutdown-seeking AIs will manipulate humans into shutting them down. W…Read more
  •  201
    This paper argues that a range of current AI systems have learned how to deceive humans. We define deception as the systematic inducement of false beliefs in the pursuit of some outcome other than the truth. We first survey empirical examples of AI deception, discussing both special-use AI systems (including Meta's CICERO) built for specific competitive situations, and general-purpose AI systems (such as large language models). Next, we detail several risks from AI deception, such as fraud, elec…Read more
  •  102
    Losing confidence in luminosity
    Noûs 55 (4): 962-991. 2021.
    A mental state is luminous if, whenever an agent is in that state, they are in a position to know that they are. Following Timothy Williamson's Knowledge and Its Limits, a wave of recent work has explored whether there are any non‐trivial luminous mental states. A version of Williamson's anti‐luminosity appeals to a safety‐theoretic principle connecting knowledge and confidence: if an agent knows p, then p is true in any nearby scenario where she has a similar level of confidence in p. However, …Read more
  •  2949
    Recent advances in natural language processing have given rise to a new kind of AI architecture: the language agent. By repeatedly calling an LLM to perform a variety of cognitive tasks, language agents are able to function autonomously to pursue goals specified in natural language and stored in a human-readable format. Because of their architecture, language agents exhibit behavior that is predictable according to the laws of folk psychology: they function as though they have desires and belief…Read more
  •  5425
    AI wellbeing
    Asian Journal of Philosophy 4 (1): 1-22. 2025.
    Under what conditions would an artificially intelligent system have wellbeing? Despite its clear bearing on the ethics of human interactions with artificial systems, this question has received little direct attention. Because all major theories of wellbeing hold that an individual’s welfare level is partially determined by their mental life, we begin by considering whether artificial systems have mental states. We show that a wide range of theories of mental states, when combined with leading th…Read more
  •  2549
    Getting Accurate about Knowledge
    with Sam Carter
    Mind 132 (525): 158-191. 2022.
    There is a large literature exploring how accuracy constrains rational degrees of belief. This paper turns to the unexplored question of how accuracy constrains knowledge. We begin by introducing a simple hypothesis: increases in the accuracy of an agent’s evidence never lead to decreases in what the agent knows. We explore various precise formulations of this principle, consider arguments in its favour, and explain how it interacts with different conceptions of evidence and accuracy. As we show…Read more
  •  1501
    Omega Knowledge Matters
    Oxford Studies in Epistemology 8 63-84. 2026.
    You omega know something when you know it, and know that you know it, and know that you know that you know it, and so on. This paper first argues that omega knowledge matters, in the sense that it is required for rational assertion, action, inquiry, and belief. The paper argues that existing accounts of omega knowledge face major challenges. One account is skeptical, claiming that we have no omega knowledge of any ordinary claims about the world. Another account embraces the KK thesis, and iden…Read more
  •  2627
    Iterated Knowledge
    Oxford University Press. 2024.
    You omega know p when you possess every iteration of knowledge of p. This book argues that omega knowledge plays a central role in philosophy. In particular, the book argues that omega knowledge is necessary for permissible assertion, action, inquiry, and belief. Although omega knowledge plays this important role, existing theories of omega knowledge are unsatisfying. One theory, KK, identifies knowledge with omega knowledge. This theory struggles to accommodate cases of inexact knowledge. The o…Read more
  •  2493
    Safety, Closure, and Extended Methods
    Journal of Philosophy 121 (1): 26-54. 2024.
    Recent research has identified a tension between the Safety principle that knowledge is belief without risk of error, and the Closure principle that knowledge is preserved by competent deduction. Timothy Williamson reconciles Safety and Closure by proposing that when an agent deduces a conclusion from some premises, the agent’s method for believing the conclusion includes their method for believing each premise. We argue that this theory is untenable because it implies problematically easy epist…Read more