Fetch and read markdown files from the internet with support for partial content extraction and intelligent section ranking...
Efficiently fetch and extract partial content from markdown files on the internet with intelligent section ranking.
scripts/fetch_md.sh <url> [output_file]
Examples:
# Print to stdout
scripts/fetch_md.sh https://bun.com/docs/runtime/sql.md
# Save to file
scripts/fetch_md.sh https://bun.com/docs/runtime/sql.md output.md
Use BM25 ranking to find most relevant sections by query:
scripts/fetch_md_structured.py <url> <query> [OPTIONS]
Examples:
# Find sections about "connection pooling"
scripts/fetch_md_structured.py https://bun.com/docs/runtime/sql.md "connection pooling"
# Get top 5 sections about "authentication" with scores
scripts/fetch_md_structured.py https://example.com/docs.md "authentication" --top 5 --show-scores
# Save results to file
scripts/fetch_md_structured.py https://example.com/docs.md "query hooks" --output result.md
When to use:
scripts/fetch_md_partial.sh <url> [OPTIONS]
Extract specific line ranges:
# Get lines 10-50
scripts/fetch_md_partial.sh https://bun.com/docs/runtime/sql.md --lines 10:50
# Get first 100 lines
scripts/fetch_md_partial.sh https://example.com/docs.md --lines 1:100
Extract specific markdown sections by header:
# Get "Installation" section
scripts/fetch_md_partial.sh https://ai-sdk.dev/docs/getting-started/expo.md --section "## Installation"
# Get "API Reference" section
scripts/fetch_md_partial.sh https://example.com/docs.md --section "# API Reference"
Note: Section extraction includes the header and all content until the next same-level header.
Find and extract lines matching a pattern:
# Search for "query" keyword
scripts/fetch_md_partial.sh https://raw.githubusercontent.com/tanstack/query/main/docs/framework/react/overview.md --pattern "query"
# Search with context (3 lines before/after)
scripts/fetch_md_partial.sh https://example.com/docs.md --pattern "API" --context 3
Add --output to save results to a file:
scripts/fetch_md_partial.sh https://example.com/docs.md --lines 1:100 --output result.md
scripts/fetch_md_partial.sh https://example.com/docs.md --section "## Setup" --output setup.md
Find relevant sections by topic (Structured Retrieval):
# Most relevant approach - uses BM25 ranking
scripts/fetch_md_structured.py <url> "your search query" --top 3
Preview documentation before full read:
# Get first 50 lines to understand structure
scripts/fetch_md_partial.sh <url> --lines 1:50
Extract specific documentation sections:
# Get installation instructions only
scripts/fetch_md_partial.sh <url> --section "## Installation"
Search for specific topics:
# Find all mentions of "authentication" with context
scripts/fetch_md_partial.sh <url> --pattern "authentication" --context 5
The fetch_md_structured.py script uses BM25 ranking to intelligently rank sections by relevance:
How it works:
Parameters:
k1=1.2: Term frequency saturation (higher = more weight to term frequency)b=0.75: Length normalization (higher = more penalty for long documents)Benefits:
https://example.com/docs.mdhttps://raw.githubusercontent.com/user/repo/main/README.mdhttps://bun.com/docs/runtime/sql.mdScripts use set -e and will exit on errors. Common issues:
# symbols)