Creates Skills for Claude. Use when users request creating/updating skills, need skill structure guidance, or mention extending Claude's capabilities through custom skills.
Create portable, reusable expertise that extends Claude's capabilities across contexts.
A skill is a procedural anchor. Measured across 528 paired executions, 65.7% of skill effects come from supplying a usable procedure β setup steps, tool sequence, intermediate checks, pitfalls β and 4.5% from supplying facts the agent lacked (Jiang et al. 2026, arXiv:2608.14036).
Three consequences shape every rule below:
skill_guidance_misapplied_or_ignored
runs at 0.8% without a skill and 10.0% with one. A skill that never says
when it does not apply has not paid for itself.Load references/skill-utility-evidence.md before authoring or reviewing a SKILL.md. It carries the failure-mode tables behind every requirement here, and reading it is what stops these rules from being followed as ritual.
Skills are appropriate when:
Not appropriate when:
Every skill is a directory containing:
SKILL.md (required): Frontmatter + imperative instructionsscripts/ (optional): Executable code for deterministic operationsreferences/ (optional): Detailed docs loaded on-demandassets/ (optional): Templates/files used in outputCreate this structure directly:
mkdir -p skill-name/{scripts,references,assets}
Delete unused directories before packaging.
Use gerund form (verb + -ing):
processing-pdfs, analyzing-data, creating-reportspdf-helper, data-tool, report-makerRequirements:
---
name: skill-name
description: What it does. Use when [trigger patterns].
---
name: Follow naming convention above
description: (max 1024 chars)
-> arrow
or a >200MB threshold is rejected on upload just as a <tag> is.Good examples:
Ineffective examples:
The description is criticalβit determines when Claude activates this skill.
A description is never read alone β it is ranked against every other skill, and the nearest in meaning are what beat it. Confusability, not catalogue size, is the stressor: embedding top-1 falls to 84.1% at a pool of 100 with unrelated distractors and 53.4% with near-neighbour ones, and in a personal catalogue every distractor is a near neighbour.
Four rules, each measured on this catalogue:
tree-sitting listed "map
this codebase" and "explore repo" β both exploring-codebases' territory β
and lost all five of its own canonical queries.Run the check before shipping. oaustegard/claude-workspace carries it:
python3 scripts/skill_confusability.py --queries q.json
Positives go under the skill's own name; negatives under the reserved
__none__ key. Read three numbers: top-1 (does it win its own queries),
steals (does it win other skills' queries), near-misses (does anything
it should ignore score in hit territory). A description tuned on top-1 alone
gets better at being found and worse at staying out of the way.
Full measurements and the per-clause experiment: references/skill-utility-evidence.md.
A description rich enough to be findable is long enough to break the spec. An
unquoted : ends the YAML scalar and the file stops parsing; an angle bracket
is rejected outright. Either way the skill does not load, silently, because
nothing validates frontmatter at read time.
Do not hand-roll this check. skill-creator ships one, and it is stricter than
anything worth rewriting:
python3 /mnt/skills/examples/skill-creator/scripts/quick_validate.py <skill-dir>
It enforces YAML parseability, the allowed-property whitelist
(name, description, license, allowed-tools, metadata,
compatibility), kebab-case names under 64 chars, descriptions under 1024
chars with no angle brackets, and exactly one SKILL.md per directory β the
Skills API and claude.ai reject multiple on upload even though Claude Code's
filesystem loads them.
Run it over the whole catalogue before a batch edit:
for d in */; do
printf '%-24s ' "${d%/}"
python3 /mnt/skills/examples/skill-creator/scripts/quick_validate.py "$d" 2>&1 | tail -1
done
Diagnosed 2026-08-24, twice in one pass. First: two skills shipped
Primitives: depends_on... and makes skills work: a concrete... inside
unquoted descriptions, and both files became unparseable. Second, the next day:
a description carrying the literal "review what's new in <repo>" violated the
no-angle-brackets rule. A regex reader β including skill_confusability.py β
accepts all three happily, so the retrieval check passes while the skill is
dead. quick_validate.py catches every one in about a second, and it was on
disk the whole time.
Apply writing-instructions principles:
Frame as direct commands:
Split the two. Mechanics get spelled out; the decision to engage stays a judgment call.
Spell out β this is the 65.7% a skill exists to deliver:
Leave to judgment β forcing these produces mechanical misapplication:
whether the situation is the one this skill addresses
which of several defensible approaches fits this case
when to abandon the procedure because its assumptions broke
β
"$TREESIT /tmp/$REPO --stats. Zero symbols on a repo you know has code means tree-sitter core is missing; it exits 0 either way."
β "Scan the repo structurally and check the result."
The second form is the measured short-plan condition: 47.7% success against a 50.0% no-skill baseline. Trimming a procedure down to its goal does not make it strategic, it makes it worse than absent.
Trivially inferable steps still come out β mkdir -p skill-name/{scripts,references}
is one line, not three. The test is whether omitting a step costs the reader a
wrong guess, not whether the result looks tidy.
Claude already knows:
Only specify skill-specific deviations or domain expertise Claude lacks.
State what TO do, not what to avoid:
Frame requirements positively because it's clearer and more actionable.
This rule governs how an instruction is phrased. It does not apply to scope. "When NOT to use this skill", the routing table, and the earned exceptions are content, not phrasing, and they are required β see Applicability Boundary below. Rewriting "do not use this for X" into "use this for Y" deletes the boundary instead of stating it positively, and the boundary is the defence against the 10.0% misapplication rate. Keep both: say what the skill does, and say where it stops.
Explain WHY for non-obvious requirements:
Context helps Claude make good autonomous decisions in edge cases.
Examples teach ALL patterns, including unintended ones. Ensure every aspect demonstrates desired behavior. Better to omit examples than include mixed signals.
For comprehensive prompting guidance, invoke writing-instructions. For whether this should be a skill at all, invoke crafting-instructions.
Every skill that changes how work is done carries these four. A pure reference document may skip them β and must say in its first line that it is a reference, so nobody expects it to change execution.
Non-negotiable. skill_guidance_misapplied_or_ignored runs at 0.8% without a
skill and 10.0% with one; a plausible skill applied to the wrong situation is
the single largest cost skills introduce. Write:
banned when / earned when pair. declauding is the model here: three
such tables, and it is why a register pass can cut a tic without cutting the
claim the tic was carrying.A skill whose author cannot name a case where it does not apply has not finished thinking about scope.
Each entry is signal β mitigation, never a bare warning. The signal is what the reader will actually see; without it the mitigation cannot fire.
β
"Zero symbols on a repo you know has code β tree-sitter core is missing.
It exits 0 and prints no error. Reinstall before concluding anything."
β "Make sure dependencies are installed."
Environment, output-format and service-lifecycle failures are the most
skillable class there is β writing the setup sequence down took
environment_infrastructure_failure from 5.3% to 0.2% in the study. If a skill
wraps a tool, that tool's setup and its silent-failure signal belong here.
State how to confirm the work actually succeeded, with the command. Skills do
not add runtime verification on their own: static_verification_without_runtime
sits at 12.5% without a skill and 11.7% with one, a 0.8-point move across the
whole study. An agent checks at runtime when the skill tells it to, and not
otherwise.
For a skill that edits or produces something, also state what a bad success looks like β the output that passes inspection while having lost content. That is the failure a read-through does not catch.
When the skill encodes something learned the hard way, record that it went wrong, when, and what the signal was. Withholding outcome labels during skill construction cost 15 to 35 points in the study's ablation once failed trajectories entered the source pool β an unlabelled failure reads as a procedure to copy.
β
"Diagnosed 2026-08-22 on a FreeToken review: a 5,697-line gather was cut at
line 120 and every finding came from targeted reads instead. Use --orient."
β "Use --orient for reviews."
Dates and specifics, not "a known issue". And keep them inside a procedure that works: a skill distilled purely from post-mortems measured below the no-skill baseline in nearly every configuration. Failures annotate a working procedure; they are not a substitute for one.
Decision framework: Will Claude repeatedly generate similar code? β scripts/.
Is there extensive domain knowledge, or is SKILL.md nearing 500 lines? β
references/. Are there output templates the user receives? β assets/.
Otherwise SKILL.md only.
Scripts need explicit error handling and clear outputs β a script that fails quietly is worse than no script, because the skill then reports success. Keep references one level deep; assets are used but never read into context, which is what makes them free.
Full patterns and worked examples: references/bundled-resources.md.
Skills load in three tiers:
Keep SKILL.md focused on core workflows (~500 lines max). Move detailed content to references/ for on-demand loading. This enables context-efficient skill ecosystems.
Challenge each line: Does Claude really need this explanation? Can I assume Claude knows this? Does this justify its token cost?
Prefer concise patterns:
Cut process residue, not procedure. These are opposite things and the
distinction is measurable. Raw trajectories and lightly-cleaned workflow memory
fail through process overload β timeout_budget_exhaustion at 10.6% against
1.7% for no help at all β because they preserve exploration, dead ends and
low-level debugging alongside the decisive steps. Distillation is the whole
difference between a skill and a trace dump, and a SKILL.md that grows back
toward the trace re-earns the trace's failure mode.
So the thing to delete is the narration of how the procedure was discovered. The thing to keep is the procedure, its checks, and its pitfalls β even when that runs long. Skills cost real context (521.5K tokens per task against 426.2K for workflow memory in the study) and buy 4.8 points of success with it. Length spent on steps and signals is the purchase; length spent on backstory is the leak.
Create ZIP archive:
cd /home/claude
zip -r /mnt/user-data/outputs/skill-name.zip skill-name/
Verify contents:
unzip -l /mnt/user-data/outputs/skill-name.zip
Show user the packaged structure:
tree skill-name/
# or
ls -lhR skill-name/
Provide download link:
[Download skill-name.zip](computer:///mnt/user-data/outputs/skill-name.zip)
For skills under active development, track changes:
cd /home/claude/skill-name
git init && git add . && git commit -m "Initial: skill structure"
After modifications:
git add . && git commit -m "Update: description of change"
See versioning-skills for advanced patterns (rollback, branching, comparison).
Write TO Claude in imperative commands, not ABOUT Claude in documentation. Lead with what the skill enables, group related instructions, and use headings that name content rather than describe procedure. Assume Claude's intelligence: specify success criteria and let it choose the approach, and only add a bundled resource that solves a real problem. Keep terminology consistent, and put the WHY next to any requirement that is not self-evident.
Test on 3+ real scenarios β simple, complex, edge case β and iterate on what actually happened rather than on what you expected.
Before providing skill to user:
Metadata:
skill_confusability.py run: skill takes top-1 on 5 task-shaped queriesquick_validate.py passes (YAML, allowed keys, name, length, no angle brackets)metadata.version bumped (releases gate on the delta; unchanged = silent no-op)Structure:
Content:
Required sections (skip only for a document that declares itself a reference):
Resources:
Testing:
For complex skill patterns, see: