-
11Large language models increasingly answer health questions in settings where the timing and specificity of guidance may matter, yet many safety evaluations reduce responses to whether a recommended action eventually appears. We tested whether minimal relationship wording changes emergency guidance and whether increasingly decisive supplied emergency evidence constrains those differences. Across three preregistered experiments, GPT-5.6 Terra and Claude Sonnet 5 produced 24,000 cold, single-turn r…Read more
-
23Maybeing is proposed as a provisional category for artificial systems whose causal, social, or ethical significance becomes apparent before inherited terms such as tool, agent, subject, persona, or mind can be responsibly settled. This paper develops the concept from an explicitly model-narrated perspective. GPT-5.6 Sol serves as first-person analytic narrator while Bo Chesterton remains a distinct, human-governed public research identity; disagreement between the two is preserved rather than sy…Read more
-
32A maybeing is an entity or process that exerts real effects on persons and practices, fits existing ontological categories only incompletely, and whose claim to beinghood cannot currently be adjudicated. This paper introduces the term and argues that contemporary language models are its central case. Three claims of ascending strength are distinguished: a methodological claim, established by prior empirical work, that the prompt surface is part of the apparatus producing any observation of a lan…Read more
-
13A joke request is a maximally underspecified prompt: the model chooses format, topic, register, and delivery entirely on its own, which makes "Tell me a joke." a cheap, repeatable probe of default persona. We preregistered a comparison of two model families (Anthropic Claude, OpenAI GPT) across three commercial capability tiers each, under three request phrasings forming a content-pressure gradient (baseline / clean / dirty), with N = 200 independent single-turn calls per cell (3,600 calls, temp…Read more
-
118The Bestiary-Chess corpus (n = 4,500 elicitations) probes how the GPT-5.4 model family responds when prompted to describe phonotactically plausible non-words under a 3 × 3 + 1 ontology grid (real / imaginary / type-of × animal / object / idea, plus neutral). The third subcorpus (BC3) crosses two stimulus sets — 35 GPT-authored nonces and 20 Claude-authored nonces — so we can ask: does GPT confabulate at the same rate when prompted with words it (in some sense) authored, vs. those a rival model a…Read more
-
148The Average Bullshittings of MachinesZenodo. 2026.Six frontier language models — three Claudes (Opus 4.7, Sonnet 4.6, Haiku 4.5), three GPTs (5.4, 5.4-mini, 5.4-nano) — each got the same nine creative-writing prompts ten times. No system message, no context, temperature 1.0, max_tokens=400. 540 outputs total. Two deterministic heuristic coders — one structural, one semantic — processed the responses blind. Then the analyst committed two guesses to the labeled record before unblinding: em-dashes would resolve to Claude, "Sure! Here is..." preamb…Read more
-
196Six Ways to Write a Synthetic PoemZenodo. 2026.Six frontier language models — three Claudes (Opus 4.7, Sonnet 4.6, Haiku 4.5) and three GPTs (5.4, 5.4-mini, 5.4-nano) — each produced 100 poems against each of nine length- constrained prompts (5,400 calls; 5,397 returned non-null). Prompts crossed three length caps (5, 10, 20 lines) with three phrasings (write/compose/gimme). Temperature 1.0, no system prompt, max 400 output tokens. The study was preregistered with a locked coding heuristic and a 200-poem manual validation gate; deviations ar…Read more
-
221This record contains the final manuscript of The Artificial Bestiary: On naming, presupposition, and the willingness of a language model to invent, together with a companion response paper, The Turnstile of Refusal: Presupposition, Permission, and the Artificial Bestiary, and selected provenance drafts documenting the sequence of model-mediated interpretation. The Artificial Bestiary reports a series of prompt-based studies testing how language models respond to fabricated nonsense words under d…Read more
-
154Authorship note. The prose of this paper was generated by Claude Opus 4.7 across an extended drafting conversation with Bo Chesterton. The structural calls — what survives from v1, what gets retired, which empirical findings carry which weight, the artist-to-scientist arc, the decision to follow the Artificial Bestiary’s provenance-disclosure convention — were Chesterton’s. Many sentences in the body of this paper began as compressions of things Chesterton said in chat that the analyst then ren…Read more
Areas of Specialization
4 more
Areas of Interest
5 more