Multi-route literature expansion + metadata normalization for evidence-first surveys.
Produces a large candidate pool (papers/papers_raw.jsonl, target ≥1200) with stable IDs and provenance, ready...
Goal: build a large, verifiable candidate pool for downstream dedupe/rank, mapping, notes, citations, and drafting.
This skill is intentionally evidence-first: if you can't reach the target size with verifiable IDs/provenance, the correct behavior is to block and ask for more exports / enable network, not to fabricate.
Always read:
references/domain_pack_overview.md — how domain packs drive topic-specific behaviorDomain packs (loaded by topic match):
assets/domain_packs/llm_agents.json — pinned classic/survey arXiv IDs for LLM agent topicsUse scripts/run.py only for:
Do not treat run.py as the place for:
queries.mdkeywords, exclude, max_results, time windowpapers/import.(csv|json|jsonl|bib)papers/arxiv_export.(csv|json|jsonl|bib)papers/imports/*.(csv|json|jsonl|bib)papers/snowball/*.(csv|json|jsonl|bib)papers/papers_raw.jsonltitle (str), authors (list[str]), year (int|""), url (str)arxiv_id and/or doiabstract (str; may be empty in offline mode)source (str) + provenance (list[dict])papers/papers_raw.csv (human scan)papers/retrieval_report.md (route counts, missing-meta stats, next actions)provenance.retrieval_policy.minimum_records, use that value; survey profiles may instead derive a stricter pool target from core_size.arxiv_id or doi, plus url).uv run python .codex/skills/literature-engineer/scripts/run.py --helpuv run python .codex/skills/literature-engineer/scripts/run.py --help.queries.md.papers/import.(csv|json|jsonl|bib), papers/arxiv_export.(csv|json|jsonl|bib), papers/imports/*.(csv|json|jsonl|bib).papers/snowball/*.(csv|json|jsonl|bib).--online and/or --snowball.ref.bib can include must-cite anchors even when keyword search misses them.r.jina.ai proxy so the pipeline can still self-boot without manual exports.0 records due to transient network errors, a simple rerun is often sufficient (the pipeline should not fabricate).Offline imports only:
papers/imports/ then run:uv run python .codex/skills/literature-engineer/scripts/run.py --workspace <workspace>Explicit offline inputs (multi-route):
uv run python .codex/skills/literature-engineer/scripts/run.py --workspace <workspace> --input path/to/a.bib --input path/to/b.jsonlOnline arXiv retrieval (needs network):
uv run python .codex/skills/literature-engineer/scripts/run.py --workspace <workspace> --onlineSnowballing (needs network unless you provide offline snowball exports):
uv run python .codex/skills/literature-engineer/scripts/run.py --workspace <workspace> --snowballSymptom:
papers/papers_raw.jsonl is below the explicit or profile-derived minimum declared by the locked Workflow.Causes:
Solutions:
papers/imports/ (multiple routes/queries).papers/snowball/.--online --snowball.Symptom:
arxiv_id and doi.Solutions:
--online to backfill arXiv IDs.