Build production pipelines for Data, ML, or AI.
Every pipeline is ETL with optional stages:
āāāāāāāāāāā āāāāāāāāāāāāā āāāāāāāāāāā āāāāāāāā ā extract ā ā ā transform ā ā ā [stage] ā ā ā load ā āāāāāāāāāāā āāāāāāāāāāāāā āāāāāāāāāāā āāāāāāāā
| Type | Stage | When |
|---|---|---|
| Data | ā | Pure ETL |
| ML | train/predict | Model training or inference |
| AI | generate | LLM enrichment |
The base never changes. Plug in what you need.
| Tool | Why |
|---|---|
hatch |
Project management, environments, builds |
uv |
Fast package installs (used by hatch) |
polars |
Fast DataFrames, lazy evaluation, no pandas baggage |
pydantic |
Config and validation, env file support |
ruff |
Linting and formatting, replaces black/isort/flake8 |
mypy |
Type checking, catches bugs before runtime |
For ML add scikit-learn. For AI add anthropic.
Install hatch (once per machine):
# Linux/WSL
sudo apt-get install pipx
pipx install hatch
# macOS
brew install hatch
Run any pipeline:
hatch run pipeline
Hatch creates the virtual environment and installs dependencies automatically on first run. No manual pip install or venv activation.
One file. config.py. Pydantic BaseSettings.
class Settings(BaseSettings):
model_config = SettingsConfigDict(env_file=".env")
# Paths
input_file: Path
output_dir: Path = Path("output")
# Stage-specific (only what you need)
batch_size: int = 100
Inherit base, compose the rest. No scattered configs. No magic strings. Every module imports this one file.
Note: .env is for local development only. Production secrets come from a secret manager (Vault, AWS Secrets Manager, etc.). The Pydantic pattern works with both - env vars get injected regardless of source.
For GCP (BigQuery, Cloud Storage): See gcp-auth SKILL.
For AWS (S3, Redshift): See aws-auth SKILL (future).
Core pipelines is platform-agnostic. Compose with auth SKILLs as needed.
project/
āāā src/
ā āāā pipeline/ # Installable package
ā āāā __init__.py
ā āāā cli.py # Click entry point
ā āāā config.py # Single source of truth
ā āāā extract.py # Read from source
ā āāā transform.py # Clean, filter, enrich
ā āāā load.py # Write to destination
ā āāā [stage].py # train.py, predict.py, generate.py
āāā data/ # Schema-conformant test data
āāā tests/
āāā pyproject.toml
āāā .env
Start with files, swap to API later. Only extract.py changes.
DEV: /data (schema-conformant files) ā transform ā load
PROD: API ā validate against same schema ā transform ā load
Why this works:
/data files must match Pydantic models)/data)Config-driven swap:
# config.py
class Settings(BaseSettings):
model_config = SettingsConfigDict(env_file=".env")
input_source: Literal["file", "api"] = "file"
data_dir: Path = Path("data")
# API settings (only needed in prod)
api_key: str | None = None
api_url: str | None = None
# extract.py
def extract(settings: Settings) -> pl.DataFrame:
if settings.input_source == "file":
return pl.scan_csv(settings.data_dir / "*.csv").collect()
else:
return fetch_from_api(settings)
The /data folder:
data/
āāā orders.csv # Matches schemas/shopify/orders.py
āāā customers.csv # Matches schemas/shopify/customers.py
āāā products.csv # Matches schemas/shopify/products.py
Files must validate against Pydantic schemas. Generate with simulators or export from real source once.
The swap is just config:
# .env.dev
INPUT_SOURCE=file
DATA_DIR=./data
# .env.prod
INPUT_SOURCE=api
API_KEY=xxx
API_URL=https://api.shopify.com
Walk in the park from dev to prod.
One module per stage. Data flows one direction. No circular imports.
Important: This is a proper package (src/pipeline/), not loose files. No PYTHONPATH = "src" hack.
Pre-commit runs before every commit:
ruff check --fix and ruff formatmypy --strictIf it passes locally, it passes in CI. Type hints everywhere.
[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"
[project]
name = "pipeline"
version = "0.1.0"
dependencies = [
"click",
"polars",
"pydantic-settings",
"rich",
]
[project.scripts]
pipeline = "pipeline.cli:main"
[project.optional-dependencies]
ml = ["scikit-learn"]
ai = ["anthropic"]
[tool.hatch.build.targets.wheel]
packages = ["src/pipeline"]
[tool.hatch.envs.default]
installer = "uv"
[tool.ruff]
line-length = 88
[tool.mypy]
strict = true
[project.scripts] creates a real command. No PYTHONPATH hack.
Run with hatch run pipeline. Hatch auto-installs on first run.
Click for parsing, Pydantic for validation:
# src/pipeline/cli.py
import click
from pipeline.config import Settings
@click.command()
def main():
"""Run the pipeline."""
settings = Settings()
# Click parses CLI, Pydantic validates
...
Usage: hatch run pipeline.
tests/
āāā unit/ # Isolated function tests
āāā integration/ # Tests hitting real services (DB, API)
āāā acceptance/ # End-to-end user scenarios
āāā evals/ # Model quality (ML accuracy, AI output quality)
āāā fixtures/ # Test data, mock responses
āāā conftest.py # Shared pytest fixtures
| Type | What | When |
|---|---|---|
| unit | Pure functions, no I/O | Every commit |
| integration | Real DB, real API | CI |
| acceptance | Full pipeline run | Pre-deploy |
| evals | Model performance, prompt quality | ML/AI only |
evals/ is ML/AI-specific - not "does code work" but "does model perform":
Run with hatch test tests/unit or hatch test tests/evals. See hatch testing docs.
When given a prompt, generate the pipeline directly from the prompt description:
Don't over-engineer. Start simple, extend when needed.
The name works on two levels:
|Built for composition from the start.
Every pipeline has two output modes:
# Run pipeline with Rich output
hatch run pipeline
# Pipe to Claude for action
hatch run pipeline | claude "extract winback list, send to klaviyo"
Separation of concerns
Unix philosophy with AI
# Traditional: tool | tool | tool
cat data.csv | grep "at-risk" | wc -l
# New: tool | AI
hatch run pipeline | claude "send winback campaign"
Uses Rich for formatted text output (tables, colors, panels) - human-readable AND pipeable:
# Rich output - pretty AND pipeable
hatch run pipeline
hatch run pipeline | claude "what do you see? what should I do?"
Rich output is human-readable AND automation-ready. Record with VHS for demos.
MCP makes it real. Claude has Klaviyo MCP, Shopify MCP, Slack MCP. "Send to Klaviyo" isn't hypothetical - Claude executes it.
Scales to 500 playbooks. Every playbook follows the same pattern. Learn once, use everywhere.
Action, not just insight. The pipeline finds at-risk customers. Claude extracts the list AND sends the winback campaign.
# View results with Rich output
hatch run pipeline
# ā Tables showing RFM segments, record with VHS
# Automation
hatch run pipeline | claude "create klaviyo segment 'Winback Q1' with at-risk customers"
# ā Claude parses output, calls Klaviyo MCP, segment created
Before: Data tool ā Human reads ā Human decides ā Human acts (hours/days)
After: hatch run pipeline | claude "do it" ā Done (seconds)
The pattern collapses the entire insight-to-action cycle:
āāāāāāāāāāāā āāāāāāāāāāā āāāāāāāāāāā
ā Pipeline ā ā ā Claude ā ā ā MCP ā
ā (sensor) ā ā (brain) ā ā (hands) ā
āāāāāāāāāāāā āāāāāāāāāāā āāāāāāāāāāā
RFM ā "at-risk" ā Klaviyo
segments customers campaign
Every playbook becomes an autonomous agent when piped to Claude. 500 playbooks = 500 specialized agents that can:
This is why the structured pipeline pattern matters. It's not just code organization - it's the foundation for AI-powered automation at scale.