Skip to main content
A skill is a folder with a SKILL.md file: frontmatter with a name and description, then instructions in plain markdown. Creating one takes no code and no tooling. Creating one that actually fires at the right moments and improves output takes exactly three disciplines, and this lesson is those three.

Discipline one: start from evidence, not imagination

The worst skills document imagined problems. The best ones are transcripts of real friction. Before writing anything, catch yourself in the act:
  1. Notice what you re-explain. If you have typed the same formatting rules, the same “always exclude test accounts”, the same tone guidance into three different conversations, that is your skill. It already exists, scattered.
  2. Write the test before the skill. Take two real requests where Claude needed your explanation, and one nearby request where the skill should stay silent. Save all three. That is your evaluation suite, and it costs five minutes.
  3. Run the baseline. Give Claude the two real requests with no skill and keep the outputs. If you cannot see what is wrong with them, you do not need the skill yet.
This ordering feels bureaucratic until the first time it saves you. A skill written from imagination tends to explain things Claude already knows and skip the one rule it actually needed.

Discipline two: the description is the API

Nothing else in your skill matters if the description does not fire. At conversation start, Claude sees only the name and description of each installed skill, and matches incoming requests against them. Vague descriptions are invisible. Broad ones are obnoxious. The forge below makes this concrete. Same skill body, three candidate descriptions, four test requests. Two of the requests should fire the skill and two should not. Put each description on the bench and run it.
The pattern that passes is always the same shape: what the skill does, in third person, followed by when to use it, with the literal words a colleague would say. “Formats quarterly business review memos in the standard template. Use when the user mentions a QBR, quarterly review, or board-prep memo.” Three trigger nouns, no adjectives wasted. Third person matters more than it looks. The description gets injected into Claude’s system prompt, where “I can help you with…” reads as someone else speaking and muddies the match. Write it like a catalog entry, not a greeting.

Discipline three: be brutally concise, then choose your freedom

Once a skill fires, its body enters the context window and competes with your actual conversation for attention. Anthropic’s authoring guidance opens with the phrase “the context window is a public good”, and its default assumption is worth tattooing somewhere: Claude is already very smart. Only write down what Claude cannot already know. Claude knows what a QBR is. It does not know that your team’s memos open with three bullets of asks before any narrative, or that revenue numbers must come from the finance dashboard and nowhere else. Skill bodies should be nearly all rules of the second kind. Every paragraph must pay rent, and the body should stay under 500 lines, moving detail into bundled files that load only when needed. The remaining choice is how much freedom to give Claude, and it depends on the terrain:
Open field

Many valid approaches. Give heuristics and trust the judgment. “Review structure, flag risky clauses, suggest tightening.”

Marked trail

A preferred pattern with room to adapt. Give a template and say where deviation is fine.

Narrow bridge

One safe way through. Give the exact steps or the exact script, and say not to improvise. “Run exactly this command. Do not add flags.”

Mismatched freedom is a quiet killer in both directions. Heuristics for a fragile migration produce disasters. A rigid script for judgment work produces work that reads like a form letter.

Build it with Claude, in two chairs

You do not write the file alone. The workflow that Anthropic’s own guidance recommends uses Claude twice, in different roles:
  1. Work through the task once, normally, explaining as you go. This conversation is Claude A.
  2. At the end, ask Claude A: “Create a skill that captures what I had to explain in this conversation. Frontmatter, description with trigger words, concise body.” Claude knows the format natively; no special setup needed.
  3. Cut what it over-explains. First drafts always include a paragraph defining something Claude already knows. Delete on sight.
  4. Test with Claude B: a fresh conversation with the skill installed, running your saved evaluation requests. Claude B has none of Claude A’s context, which is the point. If Claude B misses a rule, the rule is not prominent enough in the file.
  5. Iterate on observation, not theory. “Claude B forgot the test-account filter” is a fixable bug. Move the rule up, or sharpen it from “always filter” to “MUST filter”.
Two chairs, one loop: A drafts, B reveals what A left implicit, you referee.

The pre-flight check

Before a skill goes to your team, thirty seconds against this list:
  • Description says what and when, in third person, with real trigger words
  • Body under 500 lines, nothing Claude already knows
  • Examples are concrete pairs, not descriptions of examples
  • Any bundled files are referenced from SKILL.md, one level deep
  • Any script says clearly whether Claude should run it or read it
  • It passed your two should-fire requests and stayed silent on the near-miss
  • No dates or version references that rot (“before August, use the old API”)

Authoring guidance condensed from Anthropic’s skill authoring best practices, including the evaluation-first workflow, degrees of freedom, and the Claude A / Claude B iteration loop.