Scaffold a new AI feature powered by DSPy...
Create a project structure for building an AI-powered feature with DSPy:
$ARGUMENTS/
āāā main.py # Entry point ā run your AI feature
āāā program.py # AI logic (DSPy module)
āāā metrics.py # How to measure if the AI is working
āāā optimize.py # Make the AI better automatically
āāā evaluate.py # Test the AI's quality
āāā data.py # Training/test data loading
āāā requirements.txt # Dependencies
Ask the user:
requirements.txtdspy>=3.0
Add datasets if loading from HuggingFace. Add provider-specific packages if needed.
data.pyCreate dataset loading utilities:
import dspy
def load_data():
"""Load and prepare training/dev data.
Returns:
tuple: (trainset, devset) as lists of dspy.Example
"""
# TODO: Replace with actual data loading
examples = [
dspy.Example(input_field="...", output_field="...").with_inputs("input_field"),
]
split = int(0.8 * len(examples))
return examples[:split], examples[split:]
Adapt field names to match the user's inputs/outputs.
program.pyCreate the DSPy module. Choose the right module based on the task:
| Task type | Module | When to use |
|---|---|---|
| Simple extraction or lookup | dspy.Predict |
No reasoning needed, lowest cost |
| Needs reasoning | dspy.ChainOfThought |
Most tasks ā default choice |
| Math or computation | dspy.ProgramOfThought |
Counting, dates, calculations |
| Needs external tools | dspy.ReAct |
API calls, web search, database access |
import dspy
class MySignature(dspy.Signature):
"""Describe the task here."""
# Adapt fields to user's task
input_field: str = dspy.InputField(desc="description")
output_field: str = dspy.OutputField(desc="description")
class MyProgram(dspy.Module):
def __init__(self):
self.predict = dspy.ChainOfThought(MySignature)
def forward(self, **kwargs):
return self.predict(**kwargs)
metrics.pydef metric(example, prediction, trace=None):
"""Score how good the AI output is.
Args:
example: Expected output (ground truth)
prediction: What the AI actually produced
trace: Optional trace for optimization
Returns:
float: Score between 0 and 1
"""
# TODO: Implement task-specific metric
return prediction.output_field == example.output_field
evaluate.pyimport dspy
from dspy.evaluate import Evaluate
from program import MyProgram
from metrics import metric
from data import load_data
# Configure AI provider
lm = dspy.LM("openai/gpt-4o-mini") # or "anthropic/claude-sonnet-4-5-20250929", etc.
dspy.configure(lm=lm)
# Load data
_, devset = load_data()
# Test quality
program = MyProgram()
evaluator = Evaluate(devset=devset, metric=metric, num_threads=4, display_progress=True)
score = evaluator(program)
print(f"Score: {score}")
optimize.pyimport dspy
from program import MyProgram
from metrics import metric
from data import load_data
# Configure AI provider
lm = dspy.LM("openai/gpt-4o-mini") # or "anthropic/claude-sonnet-4-5-20250929", etc.
dspy.configure(lm=lm)
# Load data
trainset, devset = load_data()
# Automatically improve the AI prompts
program = MyProgram()
optimizer = dspy.BootstrapFewShot(metric=metric, max_bootstrapped_demos=4)
optimized = optimizer.compile(program, trainset=trainset)
# Check improvement
from dspy.evaluate import Evaluate
evaluator = Evaluate(devset=devset, metric=metric, num_threads=4, display_progress=True)
score = evaluator(optimized)
print(f"Optimized score: {score}")
# Save
optimized.save("optimized.json")
main.pyimport dspy
from program import MyProgram
# Configure AI provider
lm = dspy.LM("openai/gpt-4o-mini") # or "anthropic/claude-sonnet-4-5-20250929", etc.
dspy.configure(lm=lm)
# Load optimized version if available
program = MyProgram()
try:
program.load("optimized.json")
print("Loaded optimized program")
except FileNotFoundError:
print("Running unoptimized program")
# Run
result = program(input_field="test input")
print(result)
If the user wants to serve their AI as a web API, add these files to the project structure:
$ARGUMENTS/
āāā main.py # Entry point ā run your AI feature
āāā program.py # AI logic (DSPy module)
āāā server.py # FastAPI app ā routes and startup
āāā models.py # Pydantic request/response schemas
āāā config.py # Environment configuration
āāā metrics.py # How to measure if the AI is working
āāā optimize.py # Make the AI better automatically
āāā evaluate.py # Test the AI's quality
āāā data.py # Training/test data loading
āāā requirements.txt # Dependencies
āāā Dockerfile
āāā .env.example
server.pyfrom contextlib import asynccontextmanager
import dspy
from fastapi import FastAPI
from pydantic import BaseModel, Field
from program import MyProgram
@asynccontextmanager
async def lifespan(app: FastAPI):
lm = dspy.LM("openai/gpt-4o-mini") # or "anthropic/claude-sonnet-4-5-20250929", etc.
dspy.configure(lm=lm)
app.state.program = MyProgram()
try:
app.state.program.load("optimized.json")
except FileNotFoundError:
pass
yield
app = FastAPI(title="My AI API", lifespan=lifespan)
class QueryRequest(BaseModel):
input_field: str = Field(..., min_length=1)
class QueryResponse(BaseModel):
output_field: str
@app.post("/query", response_model=QueryResponse)
async def query(request: QueryRequest):
result = app.state.program(input_field=request.input_field)
return QueryResponse(output_field=result.output_field)
@app.get("/health")
async def health():
return {"status": "ok"}
Adapt QueryRequest/QueryResponse fields to match the user's inputs/outputs.
requirements.txtdspy>=3.0
fastapi>=0.100
uvicorn[standard]
pydantic-settings>=2.0
.env.exampleAI_MODEL_NAME=openai/gpt-4o-mini
AI_API_KEY=your-api-key-here
After generating the project, tell the user:
data.py with real training data (20+ examples). No real data yet? Use /ai-generating-data to generate synthetic training examples.evaluate.py to see how well the AI works nowoptimize.py to automatically improve qualitymain.py to use the AI.with_inputs() on Example objects. Every dspy.Example used in training must call .with_inputs("field1", "field2") to mark which fields are inputs vs expected outputs. Without this, the optimizer cannot distinguish inputs from labels and silently produces garbage demos.input_field/output_field as placeholders. Claude must rename these to match the user's actual task (e.g., email/category for email classification). Leaving generic names produces a project that runs but confuses the user.dspy.Predict is faster, cheaper, and equally accurate. Only use ChainOfThought when the task genuinely benefits from step-by-step reasoning.return prediction.answer == example.answer which returns True/False. DSPy handles booleans fine, but for weighted or partial-credit metrics, return a float between 0.0 and 1.0.Install any skill:
npx skills add lebsral/DSPy-Programming-not-prompting-LMs-skills --skill <name>
/ai-improving-accuracy/ai-generating-data/ai-serving-apis/dspy-signatures/dspy-chain-of-thought/ai-do if you do not have it ā it routes any AI problem to the right skill and is the fastest way to work: npx skills add lebsral/DSPy-Programming-not-prompting-LMs-skills --skill ai-do