Orchestrate multiple frontier LLMs (Claude, GPT-5.1, Gemini 3.0 Pro, Perplexity Sonar, Grok 4.1) for comprehensive research using LLM Council pattern with peer review and synthesis
Implements Karpathy's LLM Council pattern for superior research through parallel queries, peer review, and chairman synthesis.
Geoffrey/Claude (Native Council Member):
research.py)Python External API Orchestrator:
Use multi-model research when:
Simple Mode (Perplexity only):
Council Mode (Full council):
User: "What are the latest developments in quantum computing?"
↓
I decide: Simple query (factual, current)
↓
I call: uv run scripts/research.py --query "..." --models perplexity
↓
I read: JSON response from Perplexity
↓
I format: Markdown report with citations
↓
I save: To Obsidian Geoffrey/Research folder
↓
I return: Summary to user with Obsidian link
User: "Compare the AI strategies of OpenAI, Anthropic, and Google"
↓
I decide: Council query (comparative, complex)
↓
I call: uv run scripts/research.py --query "..." --models gpt,gemini,perplexity,grok
↓
I read: JSON with all external responses
↓
I provide: My own (Claude) research response
↓
I conduct: Peer review (each model ranks others)
↓
I request: GPT-5.1 chairman synthesis
↓
I format: Comprehensive markdown report
↓
I save: To Obsidian Geoffrey/Research folder
↓
I return: Summary with Obsidian link
All research reports saved to Obsidian include:
Citations use numeric format: [1], [2], etc.
Python Script:
cd skills/multi-model-research
uv run scripts/research.py --query "Your question" --models perplexity --output /tmp/responses.json
Config:
config.yaml - Model settings, routing rulesprompts/system_prompts.yaml - Per-model system promptsprompts/peer_review.md - Peer review templateprompts/chairman_synthesis.md - GPT-5.1 synthesis templateDependencies:
API Keys Required:
All keys configured in ~/.env file.
Simple Research:
User: "What is RAG in AI?"
I route to: Simple mode (Perplexity)
Output: Concise explanation with current examples and citations
Time: ~10 seconds
Council Research:
User: "Compare serverless vs containers for production ML workloads"
I route to: Council mode (all 4 external + me)
Process:
1. GPT-5.1: Provides comprehensive technical comparison
2. Gemini 3.0: Analyzes cost and performance trade-offs
3. Perplexity: Current industry trends and case studies
4. Grok 4.1: Developer sentiment from X/Twitter
5. Claude (me): Synthesize with nuanced analysis
6. Peer review: Each model ranks others
7. GPT-5.1 (chairman): Final synthesis
Output: Multi-perspective analysis with citations
Time: ~60 seconds
This skill implements Karpathy's LLM Council pattern released November 22, 2025.