Systematically investigate large corpus sections (100GB+) using stratified sampling, pattern recognition, and computational verification...
Purpose: Enable systematic, reproducible, token-efficient investigation of large corpus sections to inform RAG architecture design.
Methodology: Based on the proven investigation framework used to analyze the 121GB Marxists Internet Archive, achieving 95% token reduction through computational verification and stratified sampling.
Activate this skill when the user requests:
Follow this 5-phase methodology for all corpus investigations:
Goal: Understand section scope without deep reading
Tasks:
Read the index page for the section (if exists)
Run directory structure analysis:
cd /path/to/section
# Get directory tree with sizes (3 levels deep)
find . -type d -maxdepth 3 | head -50
du -h --max-depth=2 | sort -h | tail -20
# Count files by type
find . -type f -name "*.html" | wc -l
find . -type f -name "*.htm" | wc -l
find . -type f -name "*.pdf" | wc -l
# Get size distribution by subdirectory
du -sh */ 2>/dev/null | sort -h
Identify subsections and document hierarchy
Calculate total size and file counts
Output: Section overview with statistics in markdown format
Goal: Sample representative files across key dimensions
Stratification Dimensions:
Sampling Strategy:
# Sample large subsections (>1GB) - prioritize
# Read 10-15 files from largest sections
# Sample medium subsections (100MB-1GB)
# Read 5-10 files from mid-size sections
# Sample small subsections (<100MB)
# Read 3-5 files total from small sections
# Sample different time periods (if applicable)
# Find files with year patterns
find /path/to/section -name "*19[0-9][0-9]*" -o -name "*20[0-9][0-9]*" | head -20
# Sample different file depths
find /path/to/section -name "index.htm*" | head -5 # Index pages
find /path/to/section -type f -name "*.htm*" | shuf | head -10 # Random content
Target Sample Size: 15-25 files total across all dimensions
For each sampled file:
Token Optimization: Extract structure only, not full content
# When reading HTML, extract structure not content:
# - DOCTYPE and charset
# - All meta tags (name and content)
# - All heading tags (h1-h6)
# - CSS classes used
# - First paragraph only
# - Link patterns (internal, external, anchors)
# - Total word count estimate
Goal: Verify that patterns observed in samples hold across entire section
Use computational tools (grep/find), NOT exhaustive reading
Verification Commands:
# 1. Meta tag consistency
grep -r '<meta name="author"' /path/to/section | wc -l
grep -r '<meta name="description"' /path/to/section | wc -l
grep -roh '<meta name="[^"]*"' /path/to/section | sed 's/<meta name="\([^"]*\)".*/\1/' | sort | uniq -c | sort -rn
# 2. CSS class usage patterns
grep -roh 'class="[^"]*"' /path/to/section | sort | uniq -c | sort -rn | head -30
# 3. Link patterns
grep -roh 'href="[^"]*"' /path/to/section | head -100
grep -roh 'href="#[^"]*"' /path/to/section | sort | uniq -c | sort -rn | head -20
# 4. File naming conventions
find /path/to/section -type f -name "*.htm*" | sed 's/.*\///' | sort | uniq -c | sort -rn | head -30
# 5. Year/date patterns in filenames
find /path/to/section -name "*.htm*" -o -name "*.pdf" | grep -oE '(19|20)[0-9]{2}' | sort | uniq -c
# 6. DOCTYPE declarations
grep -roh '<!DOCTYPE[^>]*>' /path/to/section | sort | uniq -c
# 7. Character encoding
grep -roh 'charset=[^"]*' /path/to/section | sort | uniq -c | sort -rn
Document Confidence Levels:
Goal: Identify exceptions and unusual patterns
Sample These Outliers:
# 1. Largest files (top 5)
find /path/to/section -type f -name "*.htm*" -o -name "*.pdf" | xargs ls -lh | sort -k5 -hr | head -5
# 2. Smallest files (bottom 5)
find /path/to/section -type f -name "*.htm*" -o -name "*.pdf" | xargs ls -lh | sort -k5 -h | head -5
# 3. Files with unusual names (no standard patterns)
find /path/to/section -type f -name "*.htm*" | grep -v 'index\|chapter\|ch[0-9]\|[0-9]\{4\}'
# 4. Deepest nested files
find /path/to/section -type f -name "*.htm*" | awk '{print gsub(/\//,"/"), $0}' | sort -rn | head -10
# 5. Files without meta tags (if meta tags expected)
for file in $(find /path/to/section -name "*.htm*" | head -100); do
grep -q '<meta name' "$file" || echo "$file"
done
Read 3-5 edge case files to understand why they differ
Goal: Compile findings into actionable specification
Produce a Section Analysis Document following this structure:
# {Section Name} Analysis
**Section Path**: /absolute/path/to/section/
**Total Size**: {X}GB
**File Count**: {N} HTML, {M} PDFs
**Investigation Date**: YYYY-MM-DD
---
## 1. Executive Summary
[2-3 paragraph overview: purpose, size, key findings]
## 2. Directory Structure
[Hierarchical organization with sizes]
## 3. File Type Analysis
### HTML Files
- Count: {N}
- Naming conventions: [patterns]
- Size range: {min} - {max}
### PDF Files
- Count: {M}
- OCR quality: [assessed from samples]
- Purpose: [scanned books/periodicals/etc]
## 4. HTML Structure Patterns
### DOCTYPE and Encoding
[Common declarations]
### Meta Tag Schema
| Meta Tag | Occurrence % | Example |
|----------|--------------|---------|
| author | 95% | "Marx" |
| ... | ... | ... |
### Semantic CSS Classes
[Classes with semantic meaning]
## 5. Metadata Extraction Schema
[Define 5-layer metadata schema]
**Layer 1: File System**
- section, subsection, author (from path)
- year, chapter (from filename)
**Layer 2: HTML Meta Tags**
- title, author, description, keywords
**Layer 3: Breadcrumb Navigation**
- breadcrumb trail, category path
**Layer 4: Semantic CSS Classes**
- curator context, provenance, annotations
**Layer 5: Content-Derived**
- word_count, has_footnotes, has_images, reading_time
**Example Metadata Record**:
```json
{
"section": "...",
"author": "...",
"title": "...",
"year": 1867,
"word_count": 8500,
...
}
[What makes this section different?]
[Paragraph / Section / Article / Entry]
[Why this boundary is optimal for RAG]
[Expected token counts]
[How to preserve work β chapter β section hierarchy]
| Risk | Impact | Probability | Mitigation |
|---|---|---|---|
| ... | ... | ... | ... |
[List all files read during investigation]
[Document exact bash commands used]
# Example verification commands
grep -r '<meta name="author"' /path | wc -l
---
## Token Efficiency Tactics
**CRITICAL**: Use these tactics to achieve 95% token reduction
### 1. Computational Tools Over Reading
**DON'T**: Read 100 files to find patterns (500k tokens)
**DO**: Use grep to extract patterns + read 5 samples (25k tokens)
### 2. Strategic Sampling Over Exhaustive Coverage
**DON'T**: Read every author's archive (1M+ tokens)
**DO**: Read 3 authors (large/medium/small) + verify with find (50k tokens)
### 3. Extract Structure, Not Content
**DON'T**: Read entire documents
**DO**: Extract headings, meta tags, first paragraph only
**Implementation**:
- Read only `<head>` section for meta tags
- Extract only `<h1>-<h6>` tags for structure
- Read only first `<p>` for content sample
- Count links/images, don't read them
### 4. Aggregate Statistics Over Individual Analysis
**DON'T**: Describe each file individually
**DO**: Compute aggregate statistics and describe patterns
**Example**:
Instead of: "file1.htm has 3 meta tags" "file2.htm has 3 meta tags" "file3.htm has 2 meta tags"
Write: "95% of files have 3 meta tags (author, description, classification)" "5% are missing description tag"
### 5. Reference Examples, Don't Reproduce
**DON'T**: Include full HTML of 10 example files (50k tokens)
**DO**: Include 1-2 representative examples + reference paths (5k tokens)
### 6. Use Shell Commands for Verification
**Always prefer**:
- `grep -r` for pattern extraction
- `find` for file discovery
- `wc -l` for counting
- `sort | uniq -c` for frequency analysis
- `du -sh` for size calculations
**Over**:
- Reading files individually
- Manual counting
- Exhaustive sampling
---
## Metadata Extraction Protocol
**Extract metadata in 5 layers** for comprehensive documentation:
### Layer 1: File System Metadata
Extract from file paths using regex:
```python
# Example patterns to extract:
# - Section from /path/{section}/...
# - Author from /archive/{author}/...
# - Year from .../{year}/... or filename
# - Work slug from path structure
# - Chapter from ch##.htm filenames
Parse <meta> tags in <head>:
# Extract all meta tag names
grep -roh '<meta name="[^"]*"' /path | sed 's/<meta name="\([^"]*\)".*/\1/' | sort | uniq -c
# Extract specific meta tag
grep -roh '<meta name="author" content="[^"]*"' /path | head -20
Extract breadcrumb trails (usually <p class="breadcrumb"> or similar):
# Find breadcrumb patterns
grep -r 'class="breadcrumb"' /path | head -10
grep -r 'class="path"' /path | head -10
grep -r '<nav' /path | head -10
Identify CSS classes with semantic meaning:
# Extract all CSS classes
grep -roh 'class="[^"]*"' /path | sed 's/class="\([^"]*\)".*/\1/' | tr ' ' '\n' | sort | uniq -c | sort -rn | head -30
# Common semantic classes to look for:
# - "context" (curator annotations)
# - "information" (provenance)
# - "quoteb" (block quotes)
# - "fst" (first paragraph)
# - "title" (work titles)
Calculate from document content:
Use this strategy to select representative samples:
For a section with N total files:
1. ALWAYS read top-level index.htm (if exists)
2. Identify subsections by size:
- Large (>1GB): Sample 10-15 files
- Medium (100MB-1GB): Sample 5-10 files
- Small (<100MB): Sample 3-5 files total
3. Within each subsection, stratify by:
- File type (HTML vs PDF)
- Depth (index vs category vs content pages)
- Time period (if applicable)
4. Use random sampling within strata:
find /path -name "*.htm*" | shuf | head -10
5. Include edge cases:
- Largest file
- Smallest file
- Unusual naming pattern
- Deepest nested file
Target: 15-25 files total for sections <10GB
Target: 25-40 files total for sections >10GB
Principle: Trust but verify
After identifying a pattern from samples, verify it holds using computational tools.
# Pattern: "All files have author meta tag"
# Count total files
total=$(find /path -name "*.htm*" | wc -l)
# Count files with pattern
with_pattern=$(grep -rl '<meta name="author"' /path | wc -l)
# Calculate percentage
echo "Coverage: $with_pattern / $total = $(( 100 * with_pattern / total ))%"
For sections >10GB:
# Get random sample of 1000 files
find /path -name "*.htm*" | shuf -n 1000 > sample_files.txt
# Check pattern in sample
while read file; do
grep -q '<meta name="author"' "$file" && echo "1" || echo "0"
done < sample_files.txt | awk '{sum+=$1; count++} END {print "Coverage: " sum/count*100 "%"}'
Before completing investigation, verify:
# Find all HTML files
find /path -type f -name "*.htm*"
# Find all PDFs
find /path -type f -name "*.pdf"
# Find files modified in last 30 days
find /path -type f -mtime -30
# Find files larger than 1MB
find /path -type f -size +1M
# Find files by depth
find /path -maxdepth 2 -type f
# Extract meta tag names
grep -roh '<meta name="[^"]*"' /path | sed 's/<meta name="\([^"]*\)".*/\1/' | sort | uniq -c | sort -rn
# Extract CSS classes
grep -roh 'class="[^"]*"' /path | sort | uniq -c | sort -rn
# Extract link patterns
grep -roh 'href="[^"]*"' /path | head -100
# Extract year patterns
find /path -name "*.htm*" | grep -oE '(19|20)[0-9]{2}' | sort | uniq -c
# Extract DOCTYPE
grep -roh '<!DOCTYPE[^>]*>' /path | sort | uniq -c
# Count files by extension
find /path -type f | sed 's/.*\.//' | sort | uniq -c | sort -rn
# Size distribution by directory
du -sh /path/*/ | sort -h
# Average file size
find /path -type f -name "*.htm*" | xargs ls -l | awk '{sum+=$5; count++} END {print sum/count/1024 " KB average"}'
# Word count distribution
find /path -name "*.htm*" | head -100 | xargs wc -w | sort -n
# Random sample of 10 files
find /path -name "*.htm*" | shuf | head -10
# Stratified sample (first, middle, last)
files=$(find /path -name "*.htm*" | sort)
total=$(echo "$files" | wc -l)
echo "$files" | sed -n "1p; $(($total/2))p; ${total}p"
# Sample by size (smallest, median, largest)
find /path -name "*.htm*" | xargs ls -lh | sort -k5 -h | awk 'NR==1 || NR==total/2 || NR==total'
Here's a concrete example workflow:
USER: "Investigate the /data/archive/philosophy section"
AGENT RESPONSE:
I'll systematically investigate the philosophy section using the 5-phase methodology.
## Phase 1: Reconnaissance
[Run directory structure commands]
[Document size, file counts, hierarchy]
Section overview:
- Size: 2.3GB
- Files: 1,847 HTML, 234 PDFs
- Structure: Organized by author (45 authors)
- Time span: 1840-1950 (from filenames)
## Phase 2: Stratified Sampling
[Select 20 representative files across:]
- Large authors: Hegel (400MB), Kant (300MB), Marx (250MB)
- Medium authors: Nietzsche (150MB), Spinoza (120MB)
- Small authors: Feuerbach (40MB), Stirner (30MB)
[For each sample, extract structure:]
- Meta tags: author, title, classification
- Headings: h1 (title), h2 (chapters), h3 (sections)
- CSS classes: "title", "quoteb", "context"
- First paragraph: [sample text]
- Links: Internal chapter links, footnote anchors
## Phase 3: Pattern Verification
[Run verification commands:]
```bash
# Meta tag coverage
grep -r '<meta name="author"' /data/archive/philosophy | wc -l
# Result: 1,823 / 1,847 files = 99% coverage
# CSS class distribution
grep -roh 'class="[^"]*"' /data/archive/philosophy | sort | uniq -c | sort -rn
# Result: "quoteb" (8,234), "title" (1,847), "context" (892)
Confidence levels:
[Sample outliers:]
[Write complete section analysis document following template]
Chunking Recommendation: Section-level chunks (h3 boundaries)
RAG Integration Strategy:
Priority: HIGH (core theoretical content) Processing Complexity: MODERATE (encoding issues for German texts)
Open Questions:
---
## Reference Materials
**Full Methodology Specification**:
See `/home/user/projects/marxist-rag/docs/corpus-analysis/00-investigation-methodology-spec.md`
**Example Section Analysis**:
See `/home/user/projects/marxist-rag/docs/corpus-analysis/01-archive-section-analysis.md`
**Corpus Overview**:
See `/home/user/projects/marxist-rag/docs/corpus-analysis/00-corpus-overview.md`
---
## Success Criteria
A successful corpus investigation produces:
1. **Complete section analysis document** following the template
2. **Metadata extraction schema** ready for implementation
3. **Chunking strategy** with clear rationale
4. **RAG integration approach** with concrete steps
5. **Confidence levels** for all identified patterns
6. **Verification commands** enabling reproduction
7. **Edge case documentation** with handling recommendations
8. **Open questions** for user decision
**Token Budget**: 15,000-30,000 tokens per section investigation
**Time Estimate**: 30-60 minutes for AI agent
**Output Size**: 5,000-10,000 token specification document
---
## Investigation Principles Summary
1. **Sample strategically, verify computationally**
2. **Extract structure, not content**
3. **Document patterns, not instances**
4. **Specify confidence levels**
5. **Produce actionable specifications**
**Remember**: The goal is not to read every file, but to understand the corpus structure well enough to design an optimal RAG processing pipeline.
Use computational tools (grep, find, wc, sort, uniq) to verify patterns across thousands of files without reading them individually. This achieves 95% token reduction while maintaining investigative rigor.
---
## Notes
- This methodology was proven on the 121GB Marxists Internet Archive investigation
- Investigations following this framework are reproducible by other AI agents
- Token efficiency tactics are critical for large-scale corpus analysis
- Stratified sampling ensures representative coverage without exhaustive reading
- Computational verification provides confidence without manual checking
- Structured output enables direct implementation of RAG pipelines
**For questions or methodology improvements, see the reference documentation in `/home/user/projects/marxist-rag/docs/corpus-analysis/`**