Process large documents (200+ pages) with structure preservation, intelligent parsing, and memory-efficient handling...
β οΈ Repo Reality Check (read this first)
The real components are:
- Top-level pipeline:
LargeDocumentProcessor- Structure-aware parser:
AdvancedDocumentParser- Streaming OCR with progress:
EnhancedOCRProcessor- Chunker:
IntelligentTextChunkerβ see the intelligent-text-chunking skill.- Training data generation:
AITrainingDataGenerator- Setup helper:
scripts/setup_large_document_processing.pyThe NWT EPUB parser exposes only
get_verse(book_num, chapter, verse)(nwt_epub_parser.py) β there is noget_chapter/get_book. See the bible-epub-processing skill.Source data lives under
config/data/(NOT a top-leveldata/).Always wrap chunking calls with
protect_scripture_references/restore_scripture_referencesfromsrc/utils/scripture_parser.pywhen input may contain Bible references.
Two tightly related concerns combined here:
| File | Purpose |
|---|---|
src/utils/nwt_epub_parser.py |
EPUB parser for NWT Bible (English + Chuukese) |
scripts/extract_jwpub.py |
Extract JW publication .jwpub archives |
scripts/setup_large_document_processing.py |
One-time document pipeline setup |
output/processed_document/ |
Output directory for processed content |
python-docxPyMuPDF (import as fitz) β note: fitz==0.0.1.dev2 is NOT in requirements; use PyMuPDF onlyebooklib + NWTEpubParserfrom src.utils.nwt_epub_parser import NWTEpubParser
parser = NWTEpubParser('data/bible/nwt_E.epub')
verse_text = parser.get_verse('John', 3, 16)
chapter_verses = parser.get_chapter('Genesis', 1)
import fitz # PyMuPDF β installed as PyMuPDF, exposed as fitz
doc = fitz.open('large_document.pdf')
for page_num, page in enumerate(doc):
text = page.get_text()
# process text...
| Strategy | Use case |
|---|---|
| Semantic | AI training data β respect topic/paragraph boundaries |
| Structural | Documents with clear headings/sections |
| Fixed-size | RAG systems needing predictable chunk sizes |
| Sliding window | QA tasks needing context overlap |
# Sentence-boundary-aware chunking
def chunk_text(text: str, max_chars: int = 1024, overlap: int = 100) -> list[str]:
sentences = re.split(r'(?<=[.!?])\s+', text)
chunks, current = [], ''
for sent in sentences:
if len(current) + len(sent) > max_chars and current:
chunks.append(current.strip())
current = current[-overlap:] + ' ' + sent # overlap
else:
current += ' ' + sent
if current.strip():
chunks.append(current.strip())
return chunks
# Chuukese uses the same sentence terminators as English
SENTENCE_ENDINGS = re.compile(r'(?<=[.!?])\s+')
def detect_language(text: str) -> str:
has_accents = bool(re.search(r'[ÑéΓΓ³ΓΊ]', text))
return 'chuukese' if has_accents else 'english'
PyMuPDF==1.23.8 β PDF processing (do NOT add fitz==0.0.1.dev2)python-docx>=1.2.0ebooklib>=0.18beautifulsoup4>=4.12.0