Comprehensive toolkit for working with Google's Gemini API using the Google GenAI SDK...
Comprehensive toolkit for integrating Google's Gemini API using the google-genai Python SDK.
Use this skill to work with Gemini models for text generation, multimodal processing (images, audio, video, PDFs), structured outputs, and chat applications. The skill provides ready-to-use example scripts, API reference notes, and production-focused best practices.
uv add google-genai
from google import genai
client = genai.Client()
response = client.models.generate_content(
model="gemini-2.5-flash",
contents="Explain quantum computing simply",
)
print(response.text)
export GEMINI_API_KEY="your-key"Use for generating, summarizing, or transforming text.
Example scripts:
scripts/basic_generation.py - Simple text generationscripts/streaming_generation.py - Stream responses for better UXQuick example:
from google.genai import types
config = types.GenerateContentConfig(
temperature=0.7,
max_output_tokens=500,
)
response = client.models.generate_content(
model="gemini-2.5-flash",
contents=f"Summarize this text: {long_text}",
config=config,
)
Analyze images, extract text (OCR), or answer questions about visual content.
Example script: scripts/image_analysis.py
Quick example:
from PIL import Image
image = Image.open("photo.jpg")
response = client.models.generate_content(
model="gemini-2.5-flash",
contents=[image, "What's in this image?"],
)
Common tasks:
Extract structured JSON data from unstructured text using function calling.
Example script: scripts/function_calling.py
Quick example:
from google.genai import types
extract_person = {
"name": "extract_person",
"description": "Extract person info",
"parameters": {
"type": "object",
"properties": {
"name": {"type": "string"},
"age": {"type": "integer"},
"occupation": {"type": "string"},
},
"required": ["name"],
},
}
tools = types.Tool(function_declarations=[extract_person])
config = types.GenerateContentConfig(tools=[tools])
response = client.models.generate_content(
model="gemini-2.5-flash",
contents="John Doe is 30 and works as an engineer",
config=config,
)
function_call = response.candidates[0].content.parts[0].function_call
data = dict(function_call.args)
Build conversational applications with context management.
Example script: scripts/multimodal_chat.py
Quick example:
chat = client.chats.create(model="gemini-2.5-flash")
response = chat.send_message("Hello!")
print(response.text)
response = chat.send_message("What did I just say?")
print(response.text)
Stream long responses for better user experience.
Example script: scripts/streaming_generation.py
Quick example:
stream = client.models.generate_content_stream(
model="gemini-2.5-flash",
contents=prompt,
)
for chunk in stream:
print(chunk.text, end="", flush=True)
Choose the right model for your use case:
See the official model list for the latest stable and preview variants.
Process images, audio, video, and PDFs alongside text.
See: references/multimodal.md for comprehensive guide
Capabilities:
Enable structured outputs and tool use.
See: references/function_calling.md for detailed guide
Use cases:
Configure content safety settings per request.
Example:
from google.genai import types
config = types.GenerateContentConfig(
safety_settings=[
types.SafetySetting(
category=types.HarmCategory.HARM_CATEGORY_HATE_SPEECH,
threshold=types.HarmBlockThreshold.BLOCK_MEDIUM_AND_ABOVE,
)
]
)
response = client.models.generate_content(
model="gemini-2.5-flash",
contents=prompt,
config=config,
)
Common patterns:
import time
max_retries = 3
for attempt in range(max_retries):
try:
response = client.models.generate_content(
model="gemini-2.5-flash",
contents=prompt,
)
break
except Exception as exc:
status = getattr(exc, "status_code", None)
if status in (429, 503) and attempt < max_retries - 1:
time.sleep(2 ** attempt)
continue
raise
if response.prompt_feedback and response.prompt_feedback.block_reason:
print(f"Content blocked: {response.prompt_feedback.block_reason}")
Detailed guides available in references/ directory:
api_reference.md: Core SDK usage, models, parameters, authenticationbest_practices.md: Prompting, error handling, performance and cost tipsmultimodal.md: Image, audio, video, and PDF processing with examplesfunction_calling.md: Structured data extraction, tools, schemasReady-to-use scripts in scripts/ directory:
basic_generation.py: Simple text generationstreaming_generation.py: Streaming responsesimage_analysis.py: Image processing and analysisfunction_calling.py: Structured data extractionmultimodal_chat.py: Chat with multimodal inputsAll scripts include error handling and can be run directly or used as templates.
When to use which approach:
basic_generation.py patternstreaming_generation.py)function_calling.py)multimodal_chat.py)multimodal.md)Performance:
gemini-2.5-flash for best price-performancegemini-2.5-pro for quality and complex reasoninggemini-3-flash-preview or gemini-3-pro-preview when you want next-gen previewsmax_output_tokens to limit costsSecurity: