Plan LLM fine-tuning and evaluation experiments. Use when the user wants to design a new experiment, plan training runs, or create an experiment_summary.yaml file.
You help users plan experiments for fine-tuning and evaluating LLMs. Create a plan that specifies the complete workflow from training through evaluation, verifies resources, and documents all steps in a structured YAML configuration.
Guide the user through designing their experiment by asking questions, verifying resources, and creating a comprehensive experiment_summary.yaml file that documents the complete plan.
Quick existence check — claude.local.md must be present before designing an experiment, because every output path and SLURM default is read from it:
ls -la claude.local.md
If missing, stop and tell the user to run /ck-setup first — that skill walks them through claude.local.md.template interactively. Do not try to write a claude.local.md from this skill.
If claude.local.md is present but you suspect drift (placeholders not filled in, fields that look stale), suggest the user run /ck-setup in validate mode to get a structured health check rather than trying to validate inline here.
Run before proceeding to catch stale envs (user pulled new pins but didn't re-run pip install -e .):
python scripts/check_env.py
STALE ENV table to the user, ask whether to pip install -e . first or continue anyway.Follow the three-stage process:
param_selection.mdGuide the user through 9 interactive steps to gather all experiment parameters:
See param_selection.md for:
validation.mdBefore presenting plan to user (step 8), validate completeness:
controls.system_prompt is set — the single source for training and eval (no separate eval copy to match)See validation.md for:
experiment_generation.mdAfter user approves, create output files:
experiment_summary.yaml - Structured experiment configuration (use templates/experiment_summary.yaml)logs/design-experiment.log - Human-readable audit trail (see logging.md)Then ask about next steps (scaffold-experiment?).
See experiment_generation.md for:
logging.mdIMPORTANT: Throughout param_selection and generation, create detailed log at {experiment_dir}/logs/design-experiment.log.
What to log:
Format: Plain text with timestamped action entries
See logging.md for:
templates/Reference materials for output generation:
templates/experiment_summary.yaml - YAML schema and structure for experiment plantemplates/experiment_summary.yaml first. Do not freestyle the YAML structure — use the template schema exactly.claude.local.md for models, datasets, scratch directoriescontrols.system_prompt; it propagates to both training and eval, so parity is automatic (per-task variation: evaluation.tasks[].system_prompt)prompt defaults to controls.prompt but can vary per-task (evaluation.tasks[].prompt — eval-only prompt sweep) or per-run (runs[].parameters.prompt — a fine-tune trains on its own prompt). Author the structure, don't make the user hand-edit configs. See param_selection.md → Prompt Sweepsepochs: null, fine-tuned models use epochs: [0, 1]controls.dataset_type is required ("chat_completion" | "text_completion") — read by torchtune at training time and propagated to drive chat-template choice at eval time for every run type, including eval-onlycreate-inspect-task skill should be run firstThis skill uses the param_selection → validation → generation pattern:
| Module | Purpose |
|---|---|
| param_selection.md | 9-step interactive workflow |
| validation.md | Completeness checklist |
| experiment_generation.md | Create YAML and log files |
| logging.md | Plain text audit trail specification |
| templates/experiment_summary.yaml | YAML schema and structure |
Pattern: Three action verbs (selection, validation, generation) matching scaffold/run skills, plus cross-cutting logging and templates.
See README.md for: Complete pattern documentation and rationale.