This skill should be used when the user asks to "use Gemini Batch API", "process documents at scale", "submit a batch job", "upload files to Gemini", or needs large-scale LLM processing...
What this skill carries — grep references/ for any subject the names below miss:
!d=${CLAUDE_SKILL_DIR}; command -v skill-toc >/dev/null 2>&1 && exec skill-toc "$d"; s=$HOME/.claude/skills/plugin-utils/bin/skill-toc; [ -x "$s" ] && exec "$s" "$d"; echo "(skill-toc unavailable: references and scripts are NOT listed here — install the plugin-utils plugin, or start a new session so its bin/ reaches PATH)"
Large-scale asynchronous document processing using Google's Gemini models.
READ EXAMPLES BEFORE WRITING ANY CODE. NO EXCEPTIONS.
User asks for batch API work
↓
MANDATORY: Read examples/batch_processor.py or examples/icon_batch_vision.py
↓
Copy the pattern exactly
↓
DO NOT guess parameter names
DO NOT try wrapper types
DO NOT improvise API calls
The Batch API has non-obvious requirements that will fail silently:
dest is a config field, not a kwarg - Pass via config={"dest": "gs://..."}. Older SDKs accepted dest= directly; newer ones raise TypeError.Rationale: Previous agents wasted hours debugging API errors that the examples would have prevented. The patterns in examples/ are battle-tested production code.
dest= as a kwarg → STOP. That works on older SDKs only; the current SDK puts dest inside config={}. Read the examples.CreateBatchJobConfig object → STOP. The config is a plain dict, not a wrapper type.examples/batch_processor.py OR examples/icon_batch_vision.pyEnforcement: Writing batch API code without reading examples first violates this IRON LAW and will result in preventable errors.
# macOS: Install via nix-darwin (add to ~/nix/ configuration)
# Or if already available: gcloud --version
# Linux: Install Google Cloud SDK from official sources
curl https://sdk.cloud.google.com | bash
# Authenticate with Google Cloud Platform
gcloud auth login
# Set up Application Default Credentials for Python libraries
gcloud auth application-default login
# Enable Vertex AI API in your project
gcloud services enable aiplatform.googleapis.com
Why both auth methods?
gcloud auth login: For gsutil and gcloud CLI commandsgcloud auth application-default login: For google-generativeai Python library# Create bucket in us-central1 (required region)
gsutil mb -l us-central1 gs://your-batch-bucket
# Verify bucket location is us-central1
gsutil ls -L -b gs://your-batch-bucket | grep "Location"
See references/gcs-setup.md for complete setup guide.
Uses the Gemini File API for input. Results returned via batch_job.dest.file_name.
from google import genai
client = genai.Client() # Uses GOOGLE_API_KEY env var
# Upload JSONL to File API
uploaded = client.files.upload(
file="requests.jsonl",
config={"mime_type": "application/jsonl"}
)
# Submit batch job
job = client.batches.create(
model="gemini-2.5-flash-lite",
src=uploaded.name, # "files/..." URI
config={"display_name": "my-batch-job"}
)
# Results available at job.dest.file_name after completion
Uses GCS URIs directly. dest is a field of the config dict in the
current SDK (older SDKs accepted dest= as a kwarg — that now raises
TypeError: Batches.create() got an unexpected keyword argument 'dest').
from google import genai
# Use Vertex AI with ADC (not API key)
client = genai.Client(
vertexai=True,
project="your-project-id",
location="us-central1"
)
# Submit batch job with GCS paths.
# Current SDK signature: create(*, model, src, config)
job = client.batches.create(
model="gemini-2.5-flash-lite",
src="gs://bucket/requests.jsonl", # GCS input
config={
"display_name": "my-job",
"dest": "gs://bucket/outputs/", # GCS output (Vertex AI only!)
},
)
Verify your SDK before changing: inspect.signature(client.batches.create).
If dest is in the kwargs, the kwarg form works; otherwise use config.
Key difference: Standard API uses File API (files/...), Vertex AI uses GCS (gs://...) with dest (now a config field).
Standard API:
client.files.upload()client.batches.create(src=uploaded.name)job.dest.file_nameVertex AI:
client.batches.create(src=..., config={"dest": ...})After submitting a batch job, use Monitor instead of sleep-polling in Python:
Monitor(
description="Gemini batch job progress",
persistent=true,
timeout_ms=3600000,
command="while true; do uv run python3 -c \"import google.genai as genai; j=genai.batches.get(name='$JOB_NAME'); print(f'{j.state} | {j.name}'); exit(0 if j.state in ('JOB_STATE_SUCCEEDED','JOB_STATE_FAILED','JOB_STATE_CANCELLED') else 1)\" && break; sleep 60; done"
)
This frees the conversation to continue working while the batch runs. You get notified when the job completes or fails — no polling loop blocking your context.
Metadata must be flat primitives (no nested objects — BigQuery-backed storage). dest is a config field, not a top-level kwarg in the current SDK (Vertex AI only). Config is a plain dict (not a wrapper type).
See the Red Flags in the first Iron Law section above — the same gotchas apply here. The Key Gotchas table below summarizes all critical issues.
| Issue | Solution |
|---|---|
| Nested metadata fails | Use flat primitives or json.dumps() for complex data |
TypeError: unexpected keyword dest |
Move dest inside config={} (Vertex AI; current SDK) |
| Mixing API patterns | Standard API: File API + no dest. Vertex AI: GCS + dest |
| Auth errors with Vertex AI | Run gcloud auth application-default login |
| vertexai=True requires ADC | API key is ignored with vertexai=True |
| Missing aiplatform API | Run gcloud services enable aiplatform.googleapis.com |
| Region mismatch (Vertex) | Use us-central1 bucket only |
| Wrong URI format (Vertex) | Use gs:// not https:// |
| Invalid JSONL | Use scripts/validate_jsonl.py |
| Image batch: inline data | Use fileData.fileUri for batch, not inline |
| Duplicate IDs | Hash file content + prompt for unique IDs |
| Large PDFs fail | Split at 1,000 pages / 50MB max (the per-file limit) |
| JSON parsing fails | Use robust extraction (see gotchas.md) |
| Output not found (Vertex) | Output URI is prefix, not file path |
uploadToFileSearchStore 503 for files >10KB |
Use two-step: files.upload() then fileSearchStores.importFile() |
| File stuck in PROCESSING state | Poll files.get() until state is ACTIVE before importing |
| SDK Pager stops after first page | Use pager.hasNextPage() + pager.nextPage(), NOT for await |
Batch inlinedResponse.response.text is undefined |
Response is raw JSON, not hydrated class. Use candidates[0].content.parts[0].text |
| Store document displayName is random ID after importFile | Read bibkey from customMetadata, not displayName |
responseMimeType + tools in batch = error code 3 |
Omit responseMimeType when using tools; use prompt-based JSON instructions |
| RuntimeError: Cannot send a request, as the client has been closed | Hold ONE genai.Client for the process; an inline/per-call client is GC'd mid-request |
Vertex batch: 404 The PublisherModel <id> does not exist |
Qualify it: publishers/google/models/<id> — the bare id works only on the Standard API |
Vertex batch 404 on a model models.list() SHOWS in the region |
Listing ≠batch-servable. Fix the LOCATION, not the model: location="global" for 3.x — it accepts a us-central1 src/dest. Downgrading a tier silently changes output |
Top 3 mistakes (bolded above):
dest= as a kwarg instead of inside config={} (Vertex AI; current SDK)See references/gotchas.md for detailed solutions (now with Gotchas 10-20; 18-20 are Vertex-batch specific).
| About to | Why Wrong | Do Instead |
|---|---|---|
Write genai.Client().batches.get(...) inline, or a def client(): return genai.Client(...) factory |
The client owns an httpx pool and closes it on __del__; the request dies mid-flight with Cannot send a request, as the client has been closed. Fatal in polling loops. |
Hold a module-level singleton for the life of the process |
Run pdftotext and send the string because "text is cheaper" |
It is not: 258 tokens/page vs ~4 chars/token, and native PDF text is unbilled. Measured 22% more expensive, plus truncation and manual OCR | Send the PDF via fileData.fileUri |
Pass a bare model id to batches.create on a Vertex client |
Batch needs the publisher path; the bare id 404s even though generate_content accepts it |
publishers/google/models/<id> |
Conclude a model is unavailable — or available — from models.list() |
Listing is not a batch-availability check: us-central1 lists 3.x models that batch then 404s, because 3.x is location: global |
Submit a probe job (~20s). Fix the LOCATION; never downgrade a tier to clear a 404 — that swaps the model your evals were run on |
Reuse one generationConfig across model tiers |
3.x needs thinkingConfig pinned or it returns empty on MAX_TOKENS; 2.5 rejects the field outright |
Set thinkingConfig only for gemini-3*, per Gotcha 17/20 |
| Debug a Vertex 404 by reading docs instead of listing + submitting | The error text names the model, never the region or the availability rule — it is the same string for three different causes | Check Gotchas 19 and 20 before assuming the id is wrong |
Gemini bills a document at 258 tokens per page, and native text extracted from the PDF is not charged at all (Google's document-processing docs, verified 2026-08-31). Extracted text is billed as ordinary input at roughly 4 chars/token, so for text-heavy documents the string is the more expensive representation. Measured on a 1,313-document legal corpus — 21,393 pages, 28.2M extracted characters:
| representation | tokens |
|---|---|
| the PDFs | 5,519,394 |
| pdftotext output | 7,050,563 |
Extracting first cost 22% more and bought nothing. It also created three problems that do not exist when you send the document:
Reach for pdftotext only to triage locally (is this file a scan? how long is it?) — never as
the transport into the model.
Limits: 50 MB or 1,000 pages per file, for both inline data and Files API uploads. Batch
requests reference the document with fileData.fileUri (a GCS URI on Vertex); never inline the
bytes in a batch JSONL.
| Limit | Value |
|---|---|
| Max requests per JSONL | 10,000 |
| Max concurrent jobs | 10 |
| Max job size | 100MB |
| Job expiration | 24 hours |
Never recall a model ID or a price from training data — it is always stale. Fetch the .md.txt variants (LLM-optimized, far easier to parse than the HTML):
https://ai.google.dev/gemini-api/docs/models.md.txthttps://ai.google.dev/gemini-api/docs/pricing.md.txtReal failures this prevents (encountered 2026-08-03):
gemini-3-pro — that ID does not exist.gemini-3.1-flash-lite priced at {input 0.125, output 0.75}; the current lineup has gemini-3.5-flash-lite at {input 0.30, output 2.50} standard, {0.15, 1.25} batch.gemini-3.6-flash but no gemini-3.6-flash-lite; the newest Flash-Lite is gemini-3.5-flash-lite.Batch API pricing is 50% of standard across models.
For structured information extraction — schema-constrained JSON pulled out of documents — default to Flash or Flash-Lite. Reserve Pro for tasks needing genuine reasoning. Do not reach for Pro by default just because the task feels important.
Measured 2026-08-03 on the realpage project (SEC IPO prospectus extraction; ~16,700 input tokens/doc, ~350-600 output; identical prompt, identical 100 documents):
| model | finds the target provision | quote-verification | judge | cost/doc | 1,926-doc run |
|---|---|---|---|---|---|
gemini-3.1-pro-preview |
47% | 97.9% | 1.00 | $0.0187 | ~$36 |
gemini-3.6-flash |
62% | 95.2% | 0.85 | $0.0151 | ~$29 |
gemini-3.5-flash |
70% | 94.3% | 1.00 | $0.0149 | ~$29 |
Pro was the most conservative extractor, not the best one. It found the target provision in 47% of documents where Flash found 62-70% of the same documents. On an extraction task Pro's extra reasoning showed up as under-extraction — the failure mode that silently biases a research dataset. Scored against held-out human hand-coding (20 rows the research team coded before the pipeline existed, never having seen a machine output), all four models were identical — 85.0% exact agreement, 90% recall on real entitlements, 80% exact on those — and they failed on the same three rows. So Pro's extra reasoning bought nothing measurable, while its conservatism cost 15-23 points of detection.
Two honest caveats. The human sample was small (20 rows, 11 companies), and identical failures on identical rows says the residual errors were structural — a provision filed in an exhibit rather than the prospectus, a right held through a GP entity — not model quality. And the detection gap itself stayed unresolved: on the documents where models disagreed there was no ground truth, so which model is right on that 23-point spread was still open. Do not read this table as "Flash is more accurate"; read it as "Pro was not measurably better, and was measurably quieter."
Cost savings from Pro → Flash are smaller than people expect when the task is input-dominated. Here it was only ~20%, because Flash input is $0.75/1M against Pro's $1.00, while output — where Flash is much cheaper — was a rounding error at ~350 tokens. Flash-Lite is the only tier that cuts input price materially ($0.15/1M, ~85% saving). Work out whether the job is input- or output-dominated before assuming a Flash switch saves real money: compute mean_input_tokens * input_price vs mean_output_tokens * output_price from a Stage 2 sample (see references/scale-up-testing.md).
| Model | Use Case | Cost | Location | Thinking default |
|---|---|---|---|---|
gemini-2.5-flash-lite |
Most batch jobs | Lowest | us-central1 | OFF |
gemini-2.5-flash |
Complex extraction | Medium | us-central1 | OFF |
gemini-2.5-pro |
Highest accuracy | Highest | us-central1 | ON (cannot disable) |
gemini-3-flash-preview |
New gen, larger context | 5× flash-lite | global | HIGH (set MINIMAL!) |
gemini-3.1-flash-lite-preview |
Cheapest gen-3 | ~2× 2.5 flash-lite | global | HIGH (set MINIMAL!) |
gemini-embedding-001 |
Default for text-only (short titles, classification, retrieval over text) | Low | Standard API | n/a |
gemini-embedding-2 |
Multimodal (text+image) inputs | Low | Standard API | n/a |
text-embedding-005 |
Need Vertex Batch console visibility (legacy) | Low | us-central1 | n/a |
Critical for Gemini 3.x: Always pin thinkingConfig: {thinkingLevel: ...} in generationConfig or batch responses will silently fail with MAX_TOKENS and empty content. The level is not the same across tiers: Flash and Flash-Lite accept MINIMAL, but Pro rejects it ("Thinking level MINIMAL is not supported for this model", verified 2026-08-03) and needs LOW. Use a helper that picks the level per model — a single hardcoded constant breaks when you switch tiers. See references/gotchas.md Gotcha 17.
Critical for embedding batches: Embedding work has its own rules and failure modes — use file-based JSONL with per-row key on the Standard API; never inlined_requests (scrambles order at scale). Default to gemini-embedding-001 for text-only tasks. See references/embeddings.md and examples/embeddings_batch.py.
references/embeddings.md - NEW: Dedicated reference for embedding batches (model choice, file-based + keyed pattern, sentinel verification)references/gcs-setup.md - Complete GCS and Vertex AI setup guidereferences/gotchas.md - 20 critical production gotchas (Gemini 3.x thinking_level per tier, location='global'; embedding gotcha now lives in embeddings.md)references/best-practices.md - Idempotent IDs, state tracking, validationreferences/scale-up-testing.md - Incremental scale-up testing (LangExtract prototyping, LLM-as-judge, Vertex AI batch, gate design, input- vs output-dominated cost)references/troubleshooting.md - Common errors and debuggingreferences/vertex-ai.md - Enterprise alternative with comparisonreferences/cli-reference.md - gsutil and gcloud commandsreferences/files-api.md - Files API: upload, poll-until-ACTIVE, 48h expiry, size limitsreferences/file-search.md - File Search (managed RAG): store creation, metadata filtering, grounding metadatareferences/structured-output.md - responseJsonSchema / responseSchema: the supported schema subset, enumsexamples/icon_batch_vision.py - NEW: Batch vision analysis with Vertex AIexamples/batch_processor.py - Complete GeminiBatchProcessor classexamples/embeddings_batch.py - NEW: gemini-embedding-2 via client.batches.create_embeddings() (the only supported production path; Vertex Batch rejects this model)examples/pipeline_template.py - Customizable pipeline templatescripts/validate_jsonl.py - Validate JSONL before submissionscripts/test_single.py - Test single request before batchGemini API evolves rapidly. For API features or model names with uncertainty, verify against current documentation.