This paper extends the notion of performative prediction, or the capacity of models to influence the
phenomena they predict, to large language models (LLMs). I argue that LLMs necessarily influence what
and how humans write and speak, and that their performance depends on this influence. While certain
LLM outputs can be verified formally or against ground truth, most are evaluated against common human
practice. For such cases, I propose an interpretation of LLMs as predictive role-players, where…
Read moreThis paper extends the notion of performative prediction, or the capacity of models to influence the
phenomena they predict, to large language models (LLMs). I argue that LLMs necessarily influence what
and how humans write and speak, and that their performance depends on this influence. While certain
LLM outputs can be verified formally or against ground truth, most are evaluated against common human
practice. For such cases, I propose an interpretation of LLMs as predictive role-players, where LLM
outputs are approximations of human linguistic behaviour appropriate to social roles or recognizable
tasks. Exposure to such outputs affects action-guiding beliefs in humans, so LLM-generated “predictions”
of typical language use become interventions into how members of society communicate or act. The
epistemic habits underlying machine learning make such influence both likely and difficult to recognize,
yet emerging computer-science frameworks can be extended to study performativity in LLMs specifically.
Evidence that LLMs already change the style and content of human writing and speech, together with the
difficulty of screening AI influence out of the human-authored datasets for next-generation models,
indicates that LLMs and human behaviour have already formed a feedback loop. Recognizing
performativity in LLMs calls for a reassessment of a core idea underlying AI evaluation. For tasks with no
objective success criteria, “human-level” intelligence is not a fixed target, but something LLMs can
reshape, since the data they learn from is the very data they influence. Thus, some LLM capability gains
can be ambiguous: either LLMs are imitating humans better, or humans are becoming like LLMs, or both.
I discuss the technical and political stakes of performative LLMs and call for the development of better
methods to quantify performative prediction.