Retrieve and extract content from URLs with AI-powered summarization and structured data extraction...
Token-efficient strategies for retrieving and extracting content from URLs using exa-ai.
Use --help to see available commands and verify usage before running:
exa-ai <command> --help
MUST follow these rules when using exa-ai get-contents:
This skill inherits requirements from Common Requirements:
--livecrawl-timeout 10000 for fresh, up-to-date content instead of cached resultsEach URL counts as one piece of content. Multiple URLs increase cost linearly.
Cost strategy:
--summary instead of --text to reduce processing (and token costs)Apply these strategies:
--output-format toon for 40% fewer tokens than JSON (use when reading output directly)--summary-schema (always pipe to jq)--text-max-characters, --links, and --image-links to control output sizeIMPORTANT: Choose one approach, don't mix them:
Examples:
# β High token usage - full text
exa-ai get-contents "https://example.com" --text --livecrawl-timeout 10000
# β
Approach 1: toon format with summary (70% reduction)
exa-ai get-contents "https://example.com" --summary --livecrawl-timeout 10000 --output-format toon
# β
Approach 2: JSON + jq for summary extraction (80% reduction)
exa-ai get-contents "https://example.com" --summary --livecrawl-timeout 10000 | jq '.results[].summary'
# β
Approach 3: Schema + jq for structured extraction (85% reduction)
exa-ai get-contents "https://example.com" \
--summary \
--livecrawl-timeout 10000 \
--summary-schema '{"type":"object","properties":{"key_info":{"type":"string"}}}' | \
jq -r '.results[].summary | fromjson | .key_info'
# β Don't mix toon with jq (toon is YAML-like, not JSON)
exa-ai get-contents "https://example.com" --output-format toon | jq -r '.results'
exa-ai get-contents "https://anthropic.com" --summary --livecrawl-timeout 10000 --output-format toon
exa-ai get-contents "https://techcrunch.com" \
--summary \
--livecrawl-timeout 10000 \
--summary-query "What are the main tech news stories on this page?" | jq '.results[].summary'
exa-ai get-contents "https://www.stripe.com" \
--summary \
--livecrawl-timeout 10000 \
--summary-schema '{"type":"object","properties":{"company_name":{"type":"string"},"main_product":{"type":"string"},"target_market":{"type":"string"}}}' | jq -r '.results[].summary | fromjson'
exa-ai get-contents "https://anthropic.com,https://openai.com,https://cohere.com" \
--summary \
--livecrawl-timeout 10000 \
--output-format toon
For complete options, examples, and advanced usage, consult REFERENCE.md.
Applies to: answer, search, find-similar, get-contents
When using schema parameters (--output-schema or --summary-schema), always wrap properties in an object:
{"type":"object","properties":{"field_name":{"type":"string"}}}
DO NOT use bare properties without the object wrapper:
{"properties":{"field_name":{"type":"string"}}} // β Missing "type":"object"
Why: The Exa API requires a valid JSON Schema with an object type at the root level. Omitting this causes validation errors.
Examples:
# β
CORRECT - object wrapper included
exa-ai search "AI news" \
--summary-schema '{"type":"object","properties":{"headline":{"type":"string"}}}'
# β WRONG - missing object wrapper
exa-ai search "AI news" \
--summary-schema '{"properties":{"headline":{"type":"string"}}}'
Applies to: answer, context, search, find-similar, get-contents
toon format produces YAML-like output, not JSON. DO NOT pipe toon output to jq for parsing:
# β WRONG - toon is not JSON
exa-ai search "query" --output-format toon | jq -r '.results'
# β
CORRECT - use JSON (default) with jq
exa-ai search "query" | jq -r '.results[].title'
# β
CORRECT - use toon for direct reading only
exa-ai search "query" --output-format toon
Why: jq expects valid JSON input. toon format is designed for human readability and produces YAML-like output that jq cannot parse.
Applies to: answer, context, search, find-similar, get-contents
Pick one strategy and stick with it throughout your workflow:
Approach 1: toon only - Compact YAML-like output for direct reading
exa-ai search "query" --output-format toonApproach 2: JSON + jq - Extract specific fields programmatically
exa-ai search "query" | jq -r '.results[].title'Approach 3: Schemas + jq - Structured data extraction with validation
exa-ai search "query" --summary-schema '{...}' | jq -r '.results[].summary | fromjson'Why: Mixing approaches increases complexity and token usage. Choosing one approach optimizes for your use case.
Applies to: monitor, search (websets), research, and all skills using complex commands
When using the Bash tool with complex shell syntax, run commands directly and parse output in separate steps:
# β WRONG - nested command substitution
webset_id=$(exa-ai webset-create --search '{"query":"..."}' | jq -r '.webset_id')
# β
CORRECT - run directly, then parse
exa-ai webset-create --search '{"query":"..."}'
# Then in a follow-up command:
webset_id=$(cat output.json | jq -r '.webset_id')
Why: Complex nested $(...) command substitutions can fail unpredictably in shell environments. Running commands directly and parsing separately improves reliability and makes debugging easier.
Applies to: All skills when using complex multi-step operations
Avoid nesting multiple levels of command substitution:
# β WRONG - deeply nested
result=$(exa-ai search "$(cat query.txt | tr '\n' ' ')" --num-results $(cat config.json | jq -r '.count'))
# β
CORRECT - sequential steps
query=$(cat query.txt | tr '\n' ' ')
count=$(cat config.json | jq -r '.count')
exa-ai search "$query" --num-results $count
Why: Nested command substitutions are fragile and hard to debug when they fail. Sequential steps make each operation explicit and easier to troubleshoot.
Applies to: All skills when working with multi-step workflows
For readability and reliability, break complex operations into clear sequential steps:
# β Less maintainable - everything in one line
exa-ai webset-create --search '{"query":"startups","count":1}' | jq -r '.webset_id' | xargs -I {} exa-ai webset-search-create {} --query "AI" --behavior override
# β
More maintainable - clear steps
exa-ai webset-create --search '{"query":"startups","count":1}'
webset_id=$(jq -r '.webset_id' < output.json)
exa-ai webset-search-create $webset_id --query "AI" --behavior override
Why: Sequential steps are easier to understand, debug, and modify. Each step can be verified independently.