Guide users through building comprehensive AI evaluation strategies using Evaluation-Driven Development (EDD)
An Agent Skill for designing comprehensive AI evaluation strategies using Evaluation-Driven Development (EDD).
Eval Coach guides you through a structured 5-step framework for evaluating LLM applications:
Invoke this skill when:
Ask the user:
From answers, help define:
Guide creation of test cases:
# Test Case Template
name: descriptive_name
category: happy_path | edge_case | adversarial
inputs:
# The inputs your agent receives
query: "user query here"
context: "any context"
outputs:
# What to validate
expected_fields: [field1, field2]
should_mention: [keyword1, keyword2]
should_not_contain: [forbidden_term]
min_length: 100
max_length: 5000
Categories:
Match methods to evaluation needs:
| What to Measure | Method | Cost | When |
|---|---|---|---|
| Schema/format | Automated | Free | Always (CI) |
| Keywords present | Automated | Free | Always (CI) |
| Semantic quality | LLM-as-Judge | $0.01-0.05 | Pre-deploy |
| Relevance to input | LLM-as-Judge | $0.01-0.05 | Pre-deploy |
| Subjective quality | Human | $5-50 | Edge cases |
| Safety/compliance | Human + Automated | Varies | Always |
Integration tiers:
Tier 1: PR-Level (<5 min)
Tier 2: Pre-Deploy (15-30 min)
Tier 3: Production Monitoring (Continuous)
Track these signals:
Recommendation: Pin model versions, run weekly evals on production samples.
After completing the framework, provide:
## Evaluation Plan for [Product Name]
### Business Objectives
- Primary goal: [goal]
- Success criteria: [criteria]
### Dataset Strategy
- Total test cases: [N]
- Happy path: [N1] cases
- Edge cases: [N2] cases
- Adversarial: [N3] cases
### Evaluation Methods
| Metric | Type | Method | Threshold |
|--------|------|--------|-----------|
| ... | ... | ... | ... |
### CI/CD Integration
- PR checks: [list]
- Pre-deploy: [list]
- Monitoring: [list]
### Next Steps
1. [First action]
2. [Second action]
3. [Third action]
This skill includes starter templates in the templates/ directory:
dataset.py - LangSmith dataset creationevaluators.py - Common evaluator implementationscompare.py - Experiment comparison utilitiesSee examples/ for complete evaluation plans:
research_squad_eval.md - Multi-agent research system evaluation