Set up complete experimental infrastructure for all runs in a designed experiment...
You help users automatically set up the complete experimental infrastructure - both fine-tuning and evaluation configurations - for all runs in a designed experiment.
Orchestrate the scaffolding process by reading tool specifications from experiment_summary.yaml and launching the appropriate subagents:
scaffold-torchtune) only if at least one run has type: "fine-tuned" (see "Deciding Which Subagents to Launch" below)scaffold-inspect)This ensures the entire experiment is ready to execute from training through evaluation. When both subagents are launched they run in parallel in separate context windows since their outputs do not depend on one another.
Preparation (training) is only needed for runs that train in this experiment. Eval-only experiments ā where every run is a base model (type: "control") or a pre-existing checkpoint (type: "eval-only") ā have nothing to fine-tune, so launching scaffold-torchtune would do nothing.
Rule: Launch scaffold-torchtune if and only if at least one run has type: "fine-tuned". Always launch scaffold-inspect.
import yaml
with open(f"{experiment_dir}/experiment_summary.yaml") as f:
config = yaml.safe_load(f)
runs = config.get("runs", []) or []
finetuned_runs = [r for r in runs if isinstance(r, dict) and r.get("type") == "fine-tuned"]
needs_torchtune = len(finetuned_runs) > 0
needs_torchtune true: launch both subagents in parallel (one message, two Task calls).needs_torchtune false: launch scaffold-inspect alone. Note in the summary and log that torchtune was skipped because the experiment has no fine-tuned runs (it is eval-only).Post-decision assert (cheap insurance against a silent miss): after deciding, confirm the partition is correct ā if needs_torchtune is false, assert that finetuned_runs is genuinely empty before skipping. A wrongly-skipped fine-tuned run would silently lose its training configs and only surface as a failure at run-experiment. If the assert ever trips, do not skip ā launch scaffold-torchtune.
Current tool support:
scaffold-torchtune subagent)scaffold-inspect subagent)Future tool support: This orchestrator is designed to route to different worker subagents based on tool choices documented in experiment_summary.yaml. Future iterations may support additional frameworks.
Run before proceeding to catch stale envs (user pulled new pins but didn't re-run pip install -e .):
python scripts/check_env.py
STALE ENV table to the user, ask whether to pip install -e . first or continue anyway.data.data_generation block is present, run src/tools/experiment/prepare_data.py to materialize the declared dataset before subagents launchlogs/scaffold-experiment.logIf user runs skill without arguments:
experiment_summary.yamlIf user provides a path:
Before beginning scaffolding, perform minimal structural validation:
experiment_summary.yaml exists:
ls {experiment_dir}/experiment_summary.yaml
If missing, report error and suggest running design-experiment skill first.
DO NOT launch subagents.
experiment_summary.yaml is readable:
import yaml
with open(f"{experiment_dir}/experiment_summary.yaml") as f:
config = yaml.safe_load(f)
If unreadable or invalid YAML, report error. DO NOT launch subagents.
Note on validation division:
The subagents (scaffold-torchtune, scaffold-inspect) will perform complete validation of:
This orchestrator routes to different subagent specifications based on tool choices in experiment_summary.yaml:
Preparation tools:
torchtune ā optimizers/torchtune_agent.mdEvaluation tools:
inspect-ai ā evaluators/inspect_agent.mdAdding new tools: Create the corresponding agent file (optimizers/{tool}_agent.md or evaluators/{tool}_agent.md) and add to this mapping.
Read experiment_summary.yaml to determine which subagents to launch.
See parsing.md for:
IMPORTANT: Use the Task tool to launch worker subagents (NOT the SlashCommand tool).
Correct approach for parallel execution:
Launch both subagents in a single message with multiple Task tool calls. This runs them in parallel.
Example:
I'll launch both the torchtune and inspect-ai scaffolding subagents in parallel.
[Use Task tool with subagent_type="scaffold-torchtune"]
[Use Task tool with subagent_type="scaffold-inspect"]
Subagent prompts should:
Why this matters:
scaffold-torchtune and scaffold-inspect are launched via the Task toolIf experiment_summary.yaml contains a data.data_generation block, run the prepare_data tool before launching any subagents:
python -m cruijff_kit.tools.experiment.prepare_data {experiment_dir}
Behavior:
data.data_generation block exists or the declared dataset is generated successfully.logs/scaffold-prepare-data.log.Currently supported generators:
model_organism ā cheap, deterministic sequence datasets (src/tools/model_organisms/). See template schema for parameters.First check needs_torchtune (see "Deciding Which Subagents to Launch" above). If no run has type: "fine-tuned", skip this step entirely ā the experiment is eval-only ā and proceed to Step 2 with scaffold-inspect alone.
Otherwise, invoke the preparation subagent based on tool specification in experiment_summary.yaml.
For torchtune: See optimizers/torchtune_agent.md for:
Launch the subagent using the Task tool with the prompt template from the agent file.
Invoke the appropriate evaluation subagent based on tool specification in experiment_summary.yaml.
For inspect-ai: See evaluators/inspect_agent.md for:
Launch the subagent using the Task tool with the prompt template from the agent file.
After launching both subagents in parallel:
Processing subagent reports:
Create an orchestration log at {experiment_dir}/logs/scaffold-experiment.log that records the high-level scaffolding process.
See logging.md for:
Key principle: The orchestration log tracks coordination and timing. Detailed implementation goes in subagent logs (logs/scaffold-torchtune.log, logs/scaffold-inspect.log).
If experiment_summary.yaml not found:
design-experiment skill firstIf optimization subagent fails:
If evaluation subagent fails:
If both subagents fail:
If a subagent doesn't report back:
After completing orchestration, provide a comprehensive summary:
## Scaffold Experiment Complete
Successfully scaffolded experiment:
`/scratch/gpfs/MSALGANIK/niznik/ck-projects/capitalization/cap_4L_lora_lr_sweep_2025-10-22/`
### Fine-Tuning Configurations (scaffold-torchtune)
ā 2 runs configured successfully
**Created runs:**
- Llama-3.2-1B-Instruct_rank4/
- Llama-3.2-1B-Instruct_rank8/
**Each run contains:**
- setup_finetune.yaml (configuration)
- finetune.yaml (torchtune config)
- finetune.slurm (SLURM script)
### Evaluation Configurations (scaffold-inspect)
ā 2 evaluation cells configured successfully
**Created cells:** (per-cell layout, issue #498 ā one directory per (task, epoch))
- Llama-3.2-1B-Instruct_rank4/eval/capitalization_epoch0/
- Llama-3.2-1B-Instruct_rank8/eval/capitalization_epoch0/
**Each cell directory contains:**
- eval.yaml (per-cell evaluation configuration)
- cell.slurm (SLURM script)
- logs/ (for inspect-ai `.eval` output)
### Logs Created
- `logs/scaffold-experiment.log` - Orchestration log (this process)
- `logs/scaffold-prepare-data.log` - Data-generation details (only if `data.data_generation` block present)
- `logs/scaffold-torchtune.log` - Fine-tuning scaffolding details
- `logs/scaffold-inspect.log` - Evaluation scaffolding details
### Next Steps
**Recommended workflow:**
1. Review the generated configurations (optional)
2. Run `run-experiment` skill to execute the complete workflow:
- Fine-tuning via `run-torchtune`
- Evaluation via `run-inspect`
3. Run `explore-experiment` skill to interpret results
## Validation Before Completion
Before reporting success, verify:
- ā experiment_summary.yaml was found and read
- ā Optimization subagent was launched and reported back
- ā Evaluation subagent was launched and reported back
- ā Both subagent log files exist (i.e., logs/scaffold-torchtune.log, logs/scaffold-inspect.log)
- ā Run directories exist with expected structure (check 1-2 examples)
- ā Evaluation directories exist with expected structure (check 1-2 examples)
- ā Orchestration log was created
**Note:** You don't need to verify every file - the subagents have already done detailed verification. Just spot-check a few directories to confirm the structure is correct.
## Important Notes
### Orchestration Principles
- This skill **orchestrates** rather than implements - it launches autonomous subagents
- Each subagent maintains its own detailed log
- The orchestration log tracks high-level flow and timing
- Subagents can be run independently if needed (outside of this skill)
- Partial success is acceptable (e.g., fine-tuning configs generated but eval fails)
### Parallel Execution
- **When both subagents are needed** (the experiment has fine-tuned runs), launch them in a **single message** with multiple Task tool calls
- For an eval-only experiment (no fine-tuned runs), launch `scaffold-inspect` alone ā there is no preparation subagent to parallelize with
- Do NOT launch them sequentially in separate messages
- **Do NOT use `run_in_background: true`** ā background agents cannot surface permission prompts to the user, so all tool calls get auto-denied. Foreground parallel (multiple Task calls in one message) works correctly.
- The subagents run independently in separate context windows
- They can work simultaneously because their outputs don't depend on each other
- Wait for the launched subagent(s) to complete before proceeding to create the orchestration log
### Subagent Communication
- Each subagent receives its own prompt with specific instructions
- Subagents have full access to tools (Read, Write, Edit, Bash, etc.)
- Subagents report back once in a final message when complete
- You cannot send follow-up messages to subagents
- If a subagent needs more information, include it in the initial prompt
### Error Recovery
If scaffolding fails:
1. Check orchestration log (scaffold-experiment.log) for high-level flow
2. Check individual subagent logs (logs/scaffold-torchtune.log, logs/scaffold-inspect.log) for details
3. Fix the issue (e.g., missing inspect-ai task script, incorrect paths in claude.local.md)
4. Re-run this skill (subagents should handle existing files gracefully)
5. Or run individual subagents directly via Task tool for targeted fixes