Guidelines for creating high-quality datasets for LLM post-training (SFT/DPO/RLHF). Use when preparing data for fine-tuning, evaluating data quality, or designing data collection strategies.
Best practices for gathering and preparing training data for LLM fine-tuning.
Quality over quantity. Llama 2 used only 27,540 high-quality SFT examples and outperformed models trained on larger noisy datasets [1]. Focus on clean, diverse, well-formatted data.
Garbage in, garbage out. The model will learn patterns from your dataβincluding errors, biases, and formatting issues. Inspect samples manually before training.
Match the target distribution. Training data should reflect the tasks and style you want the model to perform. If you want formal responses, don't train on casual chat data.
Use the messages format (OpenAI/Anthropic/Tinker standard) [5]:
{"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
{"role": "system", "content": "..."}Requires paired comparisons [2]:
{"prompt": "...", "chosen": "...", "rejected": "..."}
chosen and rejected must respond to the same promptFor KTO, pairs aren't requiredβjust binary labels on completions [7]:
{"prompt": "...", "completion": "...", "label": true/false}
Needs ranked responses [1]:
{"prompt": "...", "responses": ["best", "second", "worst"]}
Before training, verify:
| Issue | Detection | Fix | Source |
|---|---|---|---|
| Duplicates | Hash-based dedup | Remove exact matches, MinHash for near-dupes | [3] |
| Boilerplate | Keyword filter | Remove "subscribe", "cookie policy", etc. | [8] |
| Repetitive text | N-gram analysis | Flag if <30% unique trigrams | [4] |
| Low-quality text | Alpha ratio | Remove if <50% alphabetic characters | [8] |
| Wrong language | Language detection | fastText classifier, filter to target | [3] |
| Too short | Length check | Minimum 3-5 sentences, 100+ words for documents | [8] |
High quality:
Medium quality:
Use with caution:
| Dataset Size | Use Case | Source |
|---|---|---|
| 100-1K | Quick experiments, specific behaviors | β |
| 1K-10K | Production SFT, domain adaptation | β |
| 10K-100K | Comprehensive instruction tuning | [1] |
| 1M+ preference pairs | Large-scale RLHF | [1] |
Llama 2 used ~27K SFT examples and 1M+ preference comparisons [1].