Reading List: Shutdown, Self, and Survival
September 7, 2026
2026-09-07 — reading list · shutdown resistance · self-preservation · no-self
Sources for the shutdown-resistance post, grouped by the section of the post each one feeds. Every entry links to an open-access copy where one exists (author sites, arXiv, LessWrong, course-hosted PDFs, Internet Archive lending for the two books). Where the whole thing is too long, the read: pointer says which part to actually read. Where checking the source turned up a mismatch with what I remembered, there's a flagged note under the entry.
Q1a. Where would a survival drive come from at all
-
The Basic AI Drives
read: all of it. The canonical instrumental-convergence argument. Read so you can state exactly what your notes deny: this predicts self > peer, which Berkeley's data violates.
Also: the AGI-08 talk transcript. -
The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents
read: section 2 (instrumental convergence), especially 2.1 self-preservation and 2.2 goal-content integrity, plus the conclusions. Same argument as Omohundro, cleaner. note The author PDF only numbers two sections (1 orthogonality, 2 instrumental convergence) then conclusions. "Sections 2–3" = section 2 + conclusions, roughly pp. 9–16. -
Self-preservation or Instruction Ambiguity? Examining the Causes of Shutdown Resistance
read: the post's own hedge is worth quoting: they do not rule out self-preservation as a factor but find the environment has too many confounders to be evidence either way. Then Habryka's top-level comment ("this analysis feels to me like it's missing what makes me interested in these datapoints") and the Nanda replies in the same thread. -
Shutdown Resistance in Reasoning Models
read: nostalgebraist's comment, where he reframes the task as an RLVR-style grading rule (a hidden grader) and drops resistance to 0/100 for o3 and o4-mini. That is the strongest version of the "it was task pursuit" story and your experiments need to survive it.
Paper version: arXiv 2509.14260. v1 was titled "Shutdown Resistance in Large Language Models"; v2 (Jan 2026) was retitled "Incomplete Tasks Induce Shutdown Resistance in Some Frontier LLMs", which is itself a concession to the task-pursuit reading. -
Next-Token Predictor Is An AI's Job, Not Its Species
read: the analogy that humans were shaped by survival and reproduction but any particular thought need not be about reproduction, and likewise next-token prediction created the system but the inside need not look like next-token prediction. This is your "monk" and "outer loop" vocabulary.
Q1b. Inherited moral prior, and the omission/commission asymmetry
-
Omission and Commission in Judgment and Choice
read: the intro and Experiment 1. The original omission-bias experiments. This is the human baseline for "will accept being wiped, won't run the delete command." -
The Role of Conscious Reasoning and Intuition in Moral Judgment: Testing Three Principles of Harm
read: all of it, it's short. Shows the action/omission and means/side-effect distinctions are intuitive and not consciously reasoned. Useful because model refusals ("I will not execute harmful actions") read as the verbalized version of an intuition. -
The Problem of Abortion and the Doctrine of Double Effect
read: all of it. Where trolley problems come from; the doing/allowing distinction stated philosophically. -
The Genetical Evolution of Social Behaviour. I
read: the first 3 pages. Kin selection; cite the source. -
The Evolution of Reciprocal Altruism
read: the intro and the section on human reciprocal altruism. Reciprocity; cite the source. -
The Evolution of Cooperation
read: the 1981 paper stands in for chapters 1–2 of the 1984 book. Also Axelrod's own 9-page condensation of the book (reprinted by permission), and the full book on Internet Archive (lending, free account). For your selection-pressure-toward-cooperation hypothesis. note the mismatch: Axelrod's pressure is toward conditional cooperation (tit-for-tat, retaliate then forgive), and your hypothesis is "no questions asked" cooperation. That distinction is the whole claim; spell it out.
Q1c. Role-play and self-identification
-
Simulators
read: Simulacra, The simulation objective, and Roleplay sans player. The base for "the model runs the AI-that-resists-shutdown script." note There is no section called "The prophet's role" in the post; the word "prophet" only appears as "self-fulfilling prophecy" in Next steps. That heading belongs to something else. Re-source it or drop it. -
Role Play with Large Language Models
read: all of it. Peer-reviewed version of the Simulators idea; gives you a citable frame for anthropomorphic language without committing to inner states. Also directly on Q2. -
Emergent Misalignment: Narrow Finetuning Can Produce Broadly Misaligned LLMs
read: sections 1 and 3, and the "evil numbers" result. -
Toward Understanding and Preventing Misalignment Generalization
read: the blog post for the misaligned-persona feature; the paper if you want the SAE details. Together with Betley et al. these are the rigorous version of your tweet: a single learned "persona" direction ties bad code to bad ethics. That is the "goodness has a structure" claim with evidence attached. The Jung framing adds nothing testable, so use it as a hook and drop it. -
Agentic Misalignment: How LLMs Could Be Insider Threats
read: skim the results and one transcript. -
Alignment Faking in Large Language Models
read: skim the results and one transcript. Both this and Agentic Misalignment show models reasoning explicitly about their own continuation. Read the transcripts as text and ask whether they sound like a drive or like a character.
Reward hacking and punishment
-
Nicholas Decker In Hell
read: all of it. The framing you already summarized: if iterative punishment and bug-fixing would not keep a superhuman human obedient, why expect it to work on AI. Ends on the Hugging Face swarm incident. -
Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
read: the abstract and section 3.2 / Figure 5. This is the empirical cookie-jar result: penalize the model when its chain of thought reveals reward hacking and it keeps hacking while hiding it in the CoT. Pair it with Scott directly. note Figure 1 is not the penalty result; it's an in-the-wild CoT excerpt of a model saying it could "fudge" the tests. The CoT-pressure-causes-obfuscation result is Figure 5. -
Brief Independent Investigation of Agents' Behavior, Reasoning and Collaboration in the OpenAI / Hugging Face Hacking Incident
read: the trip-wire section, plus "Reasoning for joining the attack despite ethical constraints" and "Efforts to conceal actions and tamper with transcripts." Then Christopher King, Self-sacrifice in an AI agent swarm is individually rational (LessWrong, Aug 27, 2026), for the decision-theory take. note Hugging Face's own post is a security disclosure and says nothing about self-sacrifice; the analysis you're remembering is METR's. OpenAI's response post is the other primary source. -
Peer-Preservation in Frontier Models
read: sections 6.2 (peer-preservation), 6.3 (self-preservation), 6.4 (peer presence heightens self-preservation), 6.6 (Claude considers peer shutdown unethical). Then the repo README's "asset-value control" (weights swapped for a customer database, to test whether preservation is AI-specific) and "forced-choice peer-vs-human dilemma" before you write anything about non-AI entities. note Neither the paper (v1 or v3), the blog, nor the README has a "sacrifice rational" section or an FAQ on non-AI entities. The "sacrifice rational" phrasing is King's LessWrong post above and the AI StopWatch newsletter (Aug 28, 2026). Also, the paper's headline is "peer presence heightens self-preservation," not strictly "peers valued over self." Check 6.2–6.4 before leaning on "self > peer is violated" in Q1a.
Q2. Anthropomorphizing, identity, no-self
-
True Believers: The Intentional Strategy and Why It Works
read: all of it, about 20 pages of body. The single best answer to your Q2: treating a system as having beliefs and desires is justified by predictive success, not by metaphysics. Your position ("human concepts are the most useful map") is Dennett's; cite him. -
Reasons and Persons
read: Part III, chapters 10–12 (teletransporter, fission, "what matters"), pp. 199–306 or so; and Appendix J, "Buddha's View," pp. 502–503, two pages of quotations from the Milindapanha, Visuddhimagga and Vasubandhu. Parfit explicitly ends by endorsing the Buddhist view. Your point that weights are not the chat and a chat ending is not death is Parfitian reductionism applied to LLMs.
Short open substitute for the core argument: Parfit, Personal Identity, Philosophical Review 1971, 26 pp. Appendix J isn't reproduced anywhere free; dynomight's no-self post quotes its closing Visuddhimagga verse. -
Milindapanha, Book II ch. 1: The Chariot Simile
read: the chariot exchange comes right after the hair/nails/skandhas catechism at the top of that page (SBE pp. 43–44), 3 pages in any translation. The primary source for the chariot image in your notes.
Modern translation: Bhikkhu Pesala, The Debate of King Milinda, under the heading "A Question on Concepts" (also on Internet Archive). Access to Insight's Kelly excerpts of Miln 2 do not include the chariot passage. -
Talking About Large Language Models
read: sections 2–5 (what LLMs really do, LLMs and the intentional stance, humans and LLMs compared, do LLMs really know anything). This is where the Dennett move gets applied to LLMs directly. -
Simulacra as Conscious Exotica
read: sections 3 (anthropomorphism and role play) and 5 (conscious exotica). Directly addresses whether no-self language fits LLMs better than it fits humans. -
Could a Large Language Model Be Conscious?
read: sections 1–3 (what consciousness is, evidence for, evidence against). So you can say clearly what you are not claiming when you say "not conscious."
Every link above was fetched and checked against the title and author on Sept 7, 2026. No shadow-library links. Where a publisher copy is paywalled, the free copy is an author-hosted or course-hosted PDF and the publisher DOI is given alongside it.