Consolidate transcripts from a channel into a single file, sorted by date (newest first), up to 800K tokens. Use when preparing transcripts for LLM context or bulk analysis.
Why? LLMs have context limits. This skill merges multiple transcripts into a single file with accurate token counting, so you can feed an entire channel's content to Claude or GPT without exceeding limits.
python scripts/consolidate_transcripts.py <channel_name>
Output: ~/Documents/YTScriber/<channel_name>/<channel_name>-consolidated.md
[!NOTE] This feature is currently a standalone script. A
ytscriber consolidateCLI command is planned for a future release.
List available channels:
ls ~/Documents/YTScriber/
| Use Case | Recommended Limit | Flag |
|---|---|---|
| Claude (200K context) | 150000 | --limit 150000 |
| GPT-4 Turbo (128K) | 100000 | --limit 100000 |
| Full archive (Claude Pro) | 800000 | (default) |
| Quick sample | 50000 | --limit 50000 |
[!TIP] The default 800K limit leaves ~200K tokens for prompts and responses when using Claude's 1M context.
python scripts/consolidate_transcripts.py <channel_name> [--limit TOKENS] [--verbose]
Examples:
# Default (800K tokens)
python scripts/consolidate_transcripts.py library-of-minds
# Custom limit for GPT-4
python scripts/consolidate_transcripts.py aws-reinvent-2025 --limit 100000
# Verbose output showing all included files
python scripts/consolidate_transcripts.py dwarkesh-patel --verbose
Check the consolidated file was created:
ls -la ~/Documents/YTScriber/<channel_name>/*-consolidated.md
| Option | Description | Default |
|---|---|---|
channel_name |
Folder name in data directory | Required |
--limit, -l |
Maximum tokens to include | 800000 |
--verbose, -v |
Show detailed file list | False |
The consolidated file includes:
| Problem | Cause | Solution |
|---|---|---|
ModuleNotFoundError: tiktoken |
tiktoken not installed | pip install tiktoken |
No transcripts found |
Empty transcripts folder | Run ytscriber download first |
FileNotFoundError |
Channel doesn't exist | Check ls ~/Documents/YTScriber/ for valid names |
| Output file is small | Few transcripts available | Use --verbose to see what was included |
| Token count seems wrong | Old tiktoken version | pip install --upgrade tiktoken |
ls ~/Documents/YTScriber/, not the YouTube channel name.ytscriber download first.cl100k_base encoding (GPT-4/Claude compatible)