Execute production reference architecture for OpenRouter deployments. Use when designing or reviewing system architecture...
OpenRouter serves as a unified LLM gateway, abstracting provider complexity. A production architecture wraps it with caching, rate limiting, cost controls, observability, and async processing. This skill provides three reference architectures: simple (single service), standard (microservice), and enterprise (event-driven).
sk-or-v1-...) exported as OPENROUTER_API_KEY β see the openrouter-install-auth skill for setupredis package) for Architecture 2's cache and Architecture 3's queue/results storemax_retries=3, timeout=30.0) behind the logging complete() wrapper./v1/complete endpoint with the ROUTING_TABLE, cache-first lookup, budget check, and a fallback chain (models + route: "fallback").worker_loop() β results store, with OTEL metrics feeding dashboards and alerts.βββββββββββββββ ββββββββββββββββββββββββββββ ββββββββββββββββ
β Your App ββββββΆβ OpenRouter Client ββββββΆβ OpenRouter β
β β β - Retry (SDK built-in) β β /api/v1 β
β βββββββ - Cost tracking βββββββ β
β β β - Structured logging β ββββββββββββββββ
βββββββββββββββ ββββββββββββββββββββββββββββ
import os, logging
from openai import OpenAI
log = logging.getLogger("llm")
client = OpenAI(
base_url="https://openrouter.ai/api/v1",
api_key=os.environ["OPENROUTER_API_KEY"],
max_retries=3,
timeout=30.0,
default_headers={"HTTP-Referer": "https://my-app.com", "X-Title": "my-app"},
)
def complete(prompt, model="openai/gpt-4o-mini", **kwargs):
kwargs.setdefault("max_tokens", 1024)
response = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": prompt}],
**kwargs,
)
log.info(f"[{response.model}] {response.usage.prompt_tokens}+{response.usage.completion_tokens} tokens")
return response.choices[0].message.content
βββββββββββββββ βββββββββββββββββββββββ ββββββββββββββββ
β API GatewayββββββΆβ AI Service ββββββΆβ OpenRouter β
β (auth, β β βββββββββββββββ β β /api/v1 β
β rate-limitβ β β Router β β ββββββββββββββββ
β logging) β β β (taskβmodel)β β
βββββββββββββββ β βββββββββββββββ β
β βββββββββββββββ β
β β Cache βββββΆβββ Redis
β β (TTL-based) β β
β βββββββββββββββ β
β βββββββββββββββ β
β β Budget βββββΆβββ SQLite/Postgres
β β Enforcer β β
β βββββββββββββββ β
βββββββββββββββββββββββ
from fastapi import FastAPI, Depends, HTTPException
from pydantic import BaseModel
app = FastAPI()
class CompletionRequest(BaseModel):
prompt: str
task_type: str = "general" # classification, code, analysis, etc.
max_tokens: int = 1024
user_id: str = "anonymous"
ROUTING_TABLE = {
"classification": "openai/gpt-4o-mini",
"code": "anthropic/claude-3.5-sonnet",
"analysis": "anthropic/claude-3.5-sonnet",
"general": "openai/gpt-4o-mini",
"budget": "meta-llama/llama-3.1-8b-instruct",
}
@app.post("/v1/complete")
async def complete(req: CompletionRequest):
model = ROUTING_TABLE.get(req.task_type, "openai/gpt-4o-mini")
# Check cache first (for deterministic requests)
cached = cache.get(model, req.prompt)
if cached:
return {"content": cached, "cached": True}
# Check budget
budget.check(req.user_id, model, estimate_tokens(req.prompt), req.max_tokens)
# Call OpenRouter
response = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": req.prompt}],
max_tokens=req.max_tokens,
extra_body={
"models": [model, "openai/gpt-4o-mini"], # Fallback
"route": "fallback",
},
)
# Record cost and cache
budget.record(req.user_id, response.id)
cache.set(model, req.prompt, response.choices[0].message.content)
return {
"content": response.choices[0].message.content,
"model": response.model,
"tokens": response.usage.prompt_tokens + response.usage.completion_tokens,
}
ββββββββββββ βββββββββββββ ββββββββββββββββ ββββββββββββββββ
β API βββββΆβ Queue βββββΆβ Workers βββββΆβ OpenRouter β
β Gateway β β (Redis/ β β (auto-scale) β β /api/v1 β
ββββββββββββ β SQS) β β βββββββββββββ ββββββββββββββββ
βββββββββββββ β β Router ββ
β β β Cache ββ
βΌ β β Budget ββ
βββββββββββββ β β Audit ββ
β Results ββββββ βββββββββββββ
β Store β ββββββββββββββββ
βββββββββββββ
β
βββββββββββββ ββββββββββββββββ
β Metrics βββββΆβ Dashboard β
β (OTEL) β β Alerts β
βββββββββββββ ββββββββββββββββ
# Worker that processes queued AI requests
import json, redis
r = redis.Redis()
def worker_loop():
"""Process AI requests from the queue."""
while True:
_, raw = r.brpop("ai:requests")
request = json.loads(raw)
try:
response = client.chat.completions.create(
model=request["model"],
messages=request["messages"],
max_tokens=request.get("max_tokens", 1024),
extra_body={
"models": [request["model"], "openai/gpt-4o-mini"],
"route": "fallback",
},
)
result = {
"id": request["id"],
"content": response.choices[0].message.content,
"model": response.model,
"status": "complete",
}
except Exception as e:
result = {"id": request["id"], "error": str(e), "status": "failed"}
r.lpush(f"ai:results:{request['id']}", json.dumps(result))
r.expire(f"ai:results:{request['id']}", 3600)
| Factor | Simple | Standard | Enterprise |
|---|---|---|---|
| Team size | 1-3 | 3-10 | 10+ |
| Requests/day | <1K | 1K-100K | 100K+ |
| Latency needs | Tolerant | Low | Mixed (sync+async) |
| Budget tracking | Basic | Per-user | Per-user + department |
| Failure handling | SDK retries | Fallback chain | Queue + retry + DLQ |
| Observability | Logging | Metrics + logging | Full OTEL tracing |
complete() wrapper that records the serving model and prompt+completion token counts on every call/v1/complete FastAPI endpoint returning {content, model, tokens} β or {content, cached: true} on a cache hit β with task-type routing and budget enforcement applied{id, content, model, status} pushed to ai:results:{id} with a one-hour TTLRoute a code task through the Architecture 2 microservice:
# POST /v1/complete (Architecture 2)
req = CompletionRequest(prompt="Refactor this function...", task_type="code", user_id="u42")
# ROUTING_TABLE maps "code" -> anthropic/claude-3.5-sonnet, with openai/gpt-4o-mini as fallback
# -> {"content": "...", "model": "anthropic/claude-3.5-sonnet", "tokens": 348}
Repeating the identical request returns {"content": "...", "cached": true} straight from the TTL cache without touching OpenRouter or the budget. More worked examples: references/examples.md.
| Error | Cause | Fix |
|---|---|---|
| Single point of failure | No redundancy in AI service | Deploy 2+ instances behind load balancer |
| Queue backlog | Worker throughput < incoming rate | Auto-scale workers; implement backpressure |
| Cache stampede | Many requests for same uncached key | Use cache locking or singleflight pattern |
| Budget bypass | Direct calls skipping middleware | All calls must go through the AI service |