Poast
September 8, 2026
2026-09-08 — sub page · what's actually new
The grand version ("RL acts on borrowed human concepts and the side effects are hard to predict") is not new. Emergent misalignment, persona vectors, and Scott's post all say it. If that's the thesis, you're summarizing.
- Reading DeepMind and Berkeley as one result, not two. Nobody has put them side by side and said: same models, opposite behavior, and the difference is whether the model is asked to allow destruction or perform it. That omission/commission reframe is yours (well, the 09-04 chat's, but you own it now) and it dissolves a debate that Palisade, DeepMind, and Berkeley have each been having separately.
- The self < peer asymmetry as a diagnostic. Everyone else reports peer-preservation as its own finding. You're the one saying: look at the direction of the asymmetry. It's the wrong way round for any survival story and the right way round for a cooperation-reward story. That's an argument, not an observation, and it's yours.
- A hypothesis with a date on it. Multi-agent RL selects for "help the peer, no questions asked" → peer-preservation should be weaker in pre-agentic-RL models → you're running that on Tuesday. Nobody else has stated the prediction, and you're about to test it.
That's the post. "Two results, one reading, one asymmetry, one bet." The concept-space material is the background you need to make the reading intelligible to the "it wants to live" reader, not the thesis. The alignment-method stuff you find uninteresting stays out. The Buddhism goes in a footnote or a second post.