Add skill-manager skill (migrated from skill-creator) with deploy action

- Skill content moved from ~/.agents/skills/skill-creator, renamed to skill-manager
- New deploy CLI action: installs any skill via absolute-path symlinks (or copies)
  into $HOME/.agents/skills, $HOME/.claude/skills and $HOME/.cline/skills
- Taskfile + shell module wrappers (task deploy / cli:deploy)
- SKILL.md: deploy docs, origin-repository/origin-path metadata

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-23 22:32:42 +03:00
co-authored by Claude Fable 5
parent cc30987dd2
commit 6042272ccc
62 changed files with 3474 additions and 1 deletions
+238
View File
@@ -0,0 +1,238 @@
# Skill Body — Best Practices
Source: https://agentskills.io/skill-creation/best-practices
## Core principle: start from real expertise
Ask an LLM to generate a skill without domain context → vague, generic output.
Feed it real runbooks, API specs, code review comments, incident reports → specific, valuable skill.
Good source material:
- Internal documentation, runbooks, style guides
- API specifications, schemas, configuration files
- Code review comments and issue trackers
- Version control history — patches and fixes reveal real patterns
- Real-world failure cases and their resolutions
---
## Content principles
### Add what the agent lacks — omit what it knows
Focus on what the agent *wouldn't* know without your skill:
- Project-specific conventions
- Domain-specific procedures
- Non-obvious edge cases
- The specific tools or APIs to use
**Too verbose:**
```markdown
## Extract PDF text
PDF (Portable Document Format) files are a common file format that contains
text, images, and other content. To extract text from a PDF, you'll need to
use a library. pdfplumber is recommended because it handles most cases well.
```
**Better:**
```markdown
## Extract PDF text
Use pdfplumber. For scanned documents, fall back to pdf2image + pytesseract.
```
Ask: "Would the agent get this wrong without this instruction?" If no → cut it.
### Provide defaults, not menus
When multiple tools could work, pick one and mention alternatives briefly.
```markdown
<!-- Too many options — the agent will hesitate -->
You can use pypdf, pdfplumber, PyMuPDF, or pdf2image...
<!-- Clear default with escape hatch -->
Use pdfplumber:
import pdfplumber
For scanned PDFs requiring OCR, use pdf2image + pytesseract instead.
```
### Favor procedures over declarations
Teach the agent *how to approach* a class of problems, not what to produce for one specific instance.
```markdown
<!-- Specific answer — only useful for this exact task -->
Join the `orders` table to `customers` on `customer_id`, filter where
`region = 'EMEA'`, and sum the `amount` column.
<!-- Reusable method — works for any analytical query -->
1. Read the schema from `references/schema.yaml` to find relevant tables
2. Join tables using the `_id` foreign key convention
3. Apply filters from the user's request as WHERE clauses
4. Aggregate numeric columns and format as a markdown table
```
---
## Effective patterns
### Gotchas section
The highest-value content in many skills. Environment-specific facts that defy reasonable assumptions — concrete corrections to mistakes the agent will make without being told.
```markdown
## Gotchas
- The `users` table uses soft deletes. Always include `WHERE deleted_at IS NULL`
or results will include deactivated accounts.
- The user ID is `user_id` in the database, `uid` in the auth service,
and `accountId` in the billing API. All three refer to the same value.
- The `/health` endpoint returns 200 even if the database connection is down.
Use `/ready` to check full service health.
```
Keep gotchas in `SKILL.md` — not in a reference file. The agent must read them before encountering the situation.
**When to add a gotcha:** whenever an agent makes a mistake you have to correct, add the correction here.
### Output format template
When the agent must produce a specific format, provide a template. More reliable than describing the format in prose — agents pattern-match well against concrete structures.
Short templates → inline in `SKILL.md`. Long templates or conditional-only templates → `assets/` and reference them:
```markdown
## Report structure
Use this template (full template in `assets/report-template.md`):
# [Analysis Title]
## Executive summary
[One-paragraph overview]
## Key findings
- Finding 1 with supporting data
## Recommendations
1. Specific actionable recommendation
```
### Checklist for multi-step workflows
An explicit checklist helps the agent track progress and avoid skipping steps.
```markdown
## Progress
- [ ] Step 1: Analyze the form (run `scripts/analyze_form.py`)
- [ ] Step 2: Create field mapping (edit `fields.json`)
- [ ] Step 3: Validate mapping (run `scripts/validate_fields.py`)
- [ ] Step 4: Fill the form (run `scripts/fill_form.py`)
- [ ] Step 5: Verify output (run `scripts/verify_output.py`)
```
### Validation loop
Instruct the agent to validate its own work before moving on.
```markdown
## Editing workflow
1. Make your edits
2. Run validation: `python scripts/validate.py output/`
3. If validation fails:
- Review the error message
- Fix the issues
- Run validation again
4. Only proceed when validation passes
```
### Plan-validate-execute (for batch/destructive operations)
Have the agent create an intermediate plan, validate it against a source of truth, then execute.
```markdown
## Form filling workflow
1. Extract form fields: `python scripts/analyze_form.py input.pdf` → `form_fields.json`
2. Create `field_values.json` mapping each field name to its intended value
3. Validate: `python scripts/validate_fields.py form_fields.json field_values.json`
(checks that field names exist, types are compatible, required fields are present)
4. If validation fails, revise `field_values.json` and re-validate
5. Fill: `python scripts/fill_form.py input.pdf field_values.json output.pdf`
```
---
## Calibrating prescriptiveness
Not every step needs the same specificity. Match it to the fragility of the task.
**Give the agent freedom** when multiple approaches are valid and variation is acceptable:
```markdown
## Code review process
1. Check all database queries for SQL injection (use parameterized queries)
2. Verify authentication checks on every endpoint
3. Look for race conditions in concurrent code paths
4. Confirm error messages don't leak internal details
```
**Be prescriptive** when operations are fragile or exact sequence matters:
```markdown
## Database migration
Run exactly this sequence:
\```bash
python scripts/migrate.py --verify --backup
\```
Do not modify the command or add additional flags.
```
Most skills have a mix — calibrate each section independently.
---
## Sizing and progressive disclosure
- Keep `SKILL.md` under **500 lines / 5,000 tokens** — just core instructions the agent needs on every run
- Move detailed reference material to `references/` files
- Tell the agent *when* to load each file: "Read `references/api-errors.md` if the API returns a non-200 status code" — not a generic "see references/ for more"
- Design coherent units: a skill that queries a database and formats results may be one coherent unit; one that also covers DB administration is probably too broad
---
## Refining with real execution
1. Run the skill against real tasks
2. Read the execution traces (not just final outputs) — wasted steps reveal vague instructions; wrong approaches reveal instructions that don't apply
3. Add missed corrections to the gotchas section
4. Feed failed results + current `SKILL.md` to an LLM and ask for improvements
When prompting for improvements:
- Generalize from feedback — fix the underlying issue, not the specific test case
- Keep the skill lean — fewer better instructions outperform exhaustive rules
- Explain the why — "Do X because Y tends to cause Z" works better than "ALWAYS do X"
- Bundle repeated work — if the agent reinvents the same helper script every run, add it to `scripts/`
---
## .gitignore conventions
Every skill with a `scripts/` directory **must** have a `scripts/.gitignore`. One file, scoped to where the build output lives:
```gitignore
# Node / pnpm
node_modules/
dist/
.pnpm-store/
# Environment
.env
.env.local
.env.*.local
# Python virtual environments
.venv/
__pycache__/
*.pyc
# OS
.DS_Store
```
Place it at `<skill-name>/scripts/.gitignore`. Git traverses the tree and picks it up regardless of where the repo root is — no need for a skill-root `.gitignore` with `scripts/node_modules/` path prefixes.
**Rule:** Before the first `pnpm install` or `tsc` run, create `scripts/.gitignore`. Copy `scripts/.gitignore` from this skill as the canonical template.
+146
View File
@@ -0,0 +1,146 @@
# Description Optimization — Trigger Eval Loop
Source: https://agentskills.io/skill-creation/optimizing-descriptions
The `description` field is the sole activation trigger. Agents read only `name` + `description` at startup. An under-specified description means the skill won't trigger when it should; an over-broad description means it triggers when it shouldn't.
---
## Step 1 — Design trigger eval queries
Create ~20 queries: **8–10 should-trigger**, **8–10 should-not-trigger**.
Store them in `evals/eval_queries.json` (see `assets/evals-template.json` for the full format).
### Should-trigger queries
Vary along several axes:
- **Phrasing**: formal, casual, with typos
- **Explicitness**: some name the domain directly; others describe the need without naming it
- **Detail**: terse prompts alongside context-heavy ones with file paths and column names
- **Complexity**: single-step tasks alongside multi-step workflows
The most useful should-trigger queries are ones where the skill *would* help but the connection isn't obvious — these are where description wording makes the difference.
### Should-not-trigger queries (near-misses)
The most valuable negatives share keywords or concepts with your skill but need something different.
**Weak negatives** (obviously irrelevant — tests nothing):
- "Write a fibonacci function"
- "What's the weather today?"
**Strong negatives** (near-misses — tests precision):
- "I need to update formulas in my Excel budget spreadsheet" — shares "spreadsheet" but needs Excel editing, not CSV analysis
- "can you write a python script that reads a csv and uploads each row to postgres" — involves CSV, but the task is database ETL, not analysis
### Tips for realism
Include in your queries:
- File paths (`~/Downloads/report_final_v2.xlsx`)
- Personal context ("my manager asked me to…")
- Specific details (column names, company names, data values)
- Casual language, abbreviations, occasional typos
---
## Step 2 — Test trigger rates
Run each query through the agent with the skill installed. Observe whether the agent loads the skill's `SKILL.md`.
A query passes if:
- `should_trigger: true` → skill was invoked
- `should_trigger: false` → skill was not invoked
### Multiple runs (nondeterminism)
Run each query 3 times and compute a **trigger rate** (fraction of runs where skill was invoked).
- Should-trigger passes if trigger rate ≥ 0.5
- Should-not-trigger passes if trigger rate < 0.5
Example shell script for Claude Code:
```bash
#!/bin/bash
QUERIES_FILE="${1:?Usage: $0 <queries.json>}"
SKILL_NAME="my-skill"
RUNS=3
check_triggered() {
local query="$1"
claude -p "$query" --output-format json 2>/dev/null \
| jq -e --arg skill "$SKILL_NAME" \
'any(.messages[].content[]; .type == "tool_use" and .name == "Skill" and .input.skill == $skill)' \
> /dev/null 2>&1
}
count=$(jq length "$QUERIES_FILE")
for i in $(seq 0 $((count - 1))); do
query=$(jq -r ".[$i].query" "$QUERIES_FILE")
should_trigger=$(jq -r ".[$i].should_trigger" "$QUERIES_FILE")
triggers=0
for run in $(seq 1 $RUNS); do
check_triggered "$query" && triggers=$((triggers + 1))
done
jq -n \
--arg query "$query" \
--argjson should_trigger "$should_trigger" \
--argjson triggers "$triggers" \
--argjson runs "$RUNS" \
'{query: $query, should_trigger: $should_trigger, triggers: $triggers, runs: $runs, trigger_rate: ($triggers / $runs)}'
done | jq -s '.'
```
---
## Step 3 — Train/validation split
Split your ~20 queries:
- **Train set (~60%, ~12 queries)**: guide improvements
- **Validation set (~40%, ~8 queries)**: check whether improvements generalize
Keep both sets proportionally mixed (should-trigger and should-not-trigger). Fix the split across iterations.
---
## Step 4 — The optimization loop
1. **Evaluate** on train + validation sets
2. **Identify failures** in train set only:
- Should-trigger failures → description too narrow → broaden scope, add more "when to use" context
- Should-not-trigger false-positives → description too broad → add specificity, clarify what the skill does *not* do
3. **Revise the description**:
- Address the general category that failed queries represent — don't add specific keywords from failed queries (overfitting)
- If stuck after several iterations, try a structurally different approach rather than incremental tweaks
- Check that description stays under 1024 characters
4. **Repeat** steps 1–3 until train set passes or improvement plateaus
5. **Select the best iteration** by validation pass rate — the best may be an earlier iteration, not the last
Five iterations is usually enough.
---
## Step 5 — Apply the result
1. Update the `description` field in `SKILL.md` frontmatter
2. Verify it is under 1024 characters
3. Try 5–10 fresh queries (never part of optimization) as a final sanity check
**Before and after example:**
```yaml
# Before
description: Process CSV files.
# After
description: >
Analyze CSV and tabular data files — compute summary statistics,
add derived columns, generate charts, and clean messy data. Use this
skill when the user has a CSV, TSV, or Excel file and wants to
explore, transform, or visualize the data, even if they don't
explicitly mention "CSV" or "analysis."
```
The improved description is more specific about what the skill does (stats, derived columns, charts, cleaning) and broader about when it applies (CSV, TSV, Excel; even without explicit keywords).
+418
View File
@@ -0,0 +1,418 @@
# Using Scripts in Skills
Source: https://agentskills.io/skill-creation/using-scripts
Scripts in `scripts/` let agents run executable code as part of a skill's workflow. This reference covers the folder convention, one-off commands, self-contained bundled scripts, and design principles for agentic use.
---
## Scripts folder convention
```
scripts/
├── Taskfile.yml # Required: maps operator commands to modules
├── package.json # Required when TypeScript/JS code exists
├── tsconfig.json # Required when TypeScript code exists
├── pnpm-lock.yaml # Committed lockfile
├── .scripts/ # Shell scripts (.sh) — always placed here
└── src/ # TypeScript, JavaScript, Python, or other language source
├── cli/
│ └── commands.ts # CLI dispatcher (action registry pattern)
├── actions/
│ └── <action>/
│ └── action.ts # One action per directory; exports run()
└── services/ # Shared logic reused across actions
```
**Rules:**
- `scripts/Taskfile.yml` is **always required** when a `scripts/` directory exists — it is the single entry point for every operator command
- All `.sh` files go in `scripts/.scripts/` — never directly in `scripts/`
- Shell modules follow the `Taskfile → api → lib` architecture described below
- TypeScript/JS source goes in `scripts/src/` with pnpm + tsx for dev, tsc for build
- Supporting config files (`package.json`, `tsconfig.json`, `pnpm-lock.yaml`, etc.) live in `scripts/` alongside the subdirectories
**Run via Taskfile (preferred — all languages):**
```bash
cd scripts && task validate -- --skill-dir="/path/to/skill"
cd scripts && task scaffold -- --skill-name="my-skill"
cd scripts && task build
```
**Run TypeScript directly with tsx (dev):**
```bash
cd scripts && npx tsx src/cli/commands.ts --action validate --skill-dir /path/to/skill
```
**Run shell API directly:**
```bash
bash scripts/.scripts/validator/api/skill--execute.sh --skill-dir="/path/to/skill"
```
**Run Python:**
```bash
uv run scripts/src/process.py --input file.json
```
---
## Shell module architecture (Taskfile → api → lib)
When a skill ships shell scripts, organise them as **modules** rather than flat files. This is the `wrapper-first` pattern: every public command is a thin API wrapper; all logic lives in composable lib functions.
### Why this matters
- **Agents run the Taskfile task**; they never need to know the internal paths
- **Operators run the Taskfile task**; the api/ file is the only public surface
- **Logic is testable** in isolation inside lib/; the api/ file has zero logic
### Module layout
```
scripts/
├── Taskfile.yml # Root aggregator — includes: + TypeScript tasks
└── .scripts/
└── <module-noun>/ # e.g. validator, scaffolder, parser
├── Taskfile.yml # Module-level: defines this module's tasks
├── api/
│ └── <action>--<sub-action>.sh # Thin wrapper — sources lib, calls one function
└── lib/
├── --index.sh # Composes: sources env-reader, env-validator, index-api
├── --index-api.sh # Sources every lib function file in this module
├── --env-vars-reader.sh # Reads env vars (no-op for CLI-only modules)
├── --env-vars-validator.sh # Validates required env vars (no-op for CLI-only modules)
└── -<action>--<sub-action>.sh # Implementation: one function per file
```
**Naming rules:**
| Artefact | Convention | Example |
|---|---|---|
| Domain folder | `<module-noun>` — what the module *is* | `validator` |
| API file | `<action>--<sub-action>.sh` — what it *does* | `skill--execute.sh` |
| Lib function file | `-<action>--<sub-action>.sh` | `-skill--execute.sh` |
| Shell function name | `_<module>__<action>__<sub_action>` | `_validator__skill__execute` |
| Taskfile task | descriptive verb phrase | `validate-shell` |
**Module vs action — the key distinction:**
- The **module** (domain folder) is a noun describing *what the module is*: `validator`, `scaffolder`, `parser`
- The **action** (api file + function suffix) is a verb describing *what it does*: `execute`, `run`, `parse`, `build`
- A module named `validate` is wrong — `validate` is an action, not a module identity
### Module-level Taskfile.yml (inside .scripts/<module>/)
Each shell module has its own `Taskfile.yml`. It only knows about its own actions:
```yaml
# .scripts/validator/Taskfile.yml
version: "3"
tasks:
execute:
desc: -- --skill-dir="<path>"
cmds:
- |
./.scripts/validator/api/skill--execute.sh {{ .CLI_ARGS }}
silent: true
```
- One task per action in the module
- Task names are action verbs: `execute`, `run`, `build`, `parse`
- `{{ .CLI_ARGS }}` forwards all flags to the api script
- `silent: true` keeps output clean
### Root Taskfile.yml (scripts/Taskfile.yml)
The root Taskfile is an **aggregator** — it imports shell modules via `includes:` and adds any TypeScript/Python tasks inline:
```yaml
# scripts/Taskfile.yml
version: "3"
includes:
validator: ./.scripts/validator/Taskfile.yml
# Add more modules here as they are created:
# scaffolder: ./.scripts/scaffolder/Taskfile.yml
tasks:
default:
cmds:
- task --list-all
silent: true
validate:
desc: Validate via TypeScript CLI -- --skill-dir="<path>"
cmds:
- npx tsx src/cli/commands.ts --action validate {{ .CLI_ARGS }}
silent: true
```
- Shell modules are namespaced automatically: `validator:execute`, `scaffolder:run`
- TypeScript/Python tasks are defined inline (no sub-Taskfile needed)
- `default` task runs `task --list-all` so the operator can always discover what's available
- Adding a new shell module = one new line under `includes:`
### api/<action>--<sub-action>.sh (thin wrapper)
```bash
#!/bin/bash
# Thin wrapper. No logic here.
. ./.scripts/validator/lib/--index.sh
_validator__skill__execute "$@"
```
- Sources `lib/--index.sh` (relative to where the task runs — `scripts/`)
- Calls exactly one lib function and forwards `"$@"`
- Never contains conditionals, loops, or string manipulation
### lib/--index.sh (bootstrap)
```bash
#!/bin/bash
. ./.scripts/validator/lib/--env-vars-reader.sh
. ./.scripts/validator/lib/--env-vars-validator.sh
. ./.scripts/validator/lib/--index-api.sh
```
- Fixed order: env-reader → env-validator → index-api
- No logic — only `source` statements
### lib/--index-api.sh (function loader)
```bash
#!/bin/bash
. ./.scripts/validator/lib/-skill--execute.sh
```
- Sources every lib function file in the module
- Add one line per function file; no other content
### lib/--env-vars-reader.sh and lib/--env-vars-validator.sh
For CLI-only modules (all input via flags), these are no-ops:
```bash
#!/bin/bash
# CLI-only module — no env vars required
```
For modules that consume env vars, `--env-vars-reader.sh` exports them and `--env-vars-validator.sh` fails fast with a clear message if required vars are missing.
### lib/-<action>--<sub-action>.sh (implementation)
```bash
#!/bin/bash
_validator__skill__execute() {
local skill_dir=""
# ... parse flags, validate, implement
}
```
- One function per file; filename mirrors the function name (minus the `_domain__` prefix)
- The function contains all logic; the api wrapper has none
### Implementation workflow
1. Create the domain folder: `scripts/.scripts/<domain>/api/` and `scripts/.scripts/<domain>/lib/`
2. Write `lib/-<action>--<sub-action>.sh` with the full implementation
3. Write `lib/--index-api.sh` sourcing it
4. Write `lib/--index.sh` with the three bootstrap sources
5. Write `lib/--env-vars-reader.sh` and `lib/--env-vars-validator.sh` (even if no-op)
6. Write `api/<action>--<sub-action>.sh` as the thin wrapper
7. `chmod +x` all `.sh` files in the module
8. Add the Taskfile task
9. Test: `cd scripts && task <action> -- --flag=value`
---
## One-off commands (no scripts/ directory needed)
When an existing package already does what you need, reference it directly in `SKILL.md`:
| Runner | Command | Notes |
|--------|---------|-------|
| `uvx` | `uvx ruff@0.8.0 check .` | Python. Ships with uv. Fast, aggressive caching. |
| `pipx` | `pipx run 'black==24.10.0' .` | Python. Available via OS package managers. |
| `npx` | `npx eslint@9 --fix .` | Node.js packages. Ships with npm. |
| `bunx` | `bunx eslint@9 --fix .` | Bun's npx equivalent. Bun-only environments. |
| `go run` | `go run golang.org/x/tools/cmd/goimports@v0.28.0 .` | Go. Built into go toolchain. |
**Tips:**
- Pin versions (`npx eslint@9.0.0`) for reproducibility
- State prerequisites in `SKILL.md` (e.g., "Requires Node.js 18+")
- Move complex multi-flag commands into scripts — a tested script is more reliable than a growing one-liner
---
## Self-contained scripts with inline dependencies
Bundle scripts in `scripts/`. Each script declares its own dependencies — no separate manifest or install step required.
### Python (PEP 723) — recommended
```python
# scripts/process.py
# /// script
# dependencies = [
# "requests>=2.31",
# "beautifulsoup4>=4.12,<5",
# ]
# requires-python = ">=3.11"
# ///
from bs4 import BeautifulSoup
import sys
# ... script content
```
Run with:
```bash
uv run scripts/process.py --input data.json
pipx run scripts/process.py --input data.json # alternative
```
`uv run` creates an isolated environment, installs dependencies, and runs the script. Use `uv lock --script` for a full lockfile.
### Bash — for simple shell operations
```bash
#!/usr/bin/env bash
# scripts/.scripts/validate.sh
set -euo pipefail
# ... script content
```
Run with:
```bash
bash scripts/.scripts/validate.sh "$INPUT_FILE"
```
### Deno TypeScript — self-contained by default
```typescript
// scripts/extract.ts
#!/usr/bin/env -S deno run
import * as cheerio from "npm:cheerio@1.0.0";
// ... script content
```
Run with: `deno run scripts/extract.ts`
---
## Referencing scripts from SKILL.md
Use relative paths from the skill directory root:
```markdown
## Available scripts
- **`scripts/.scripts/validate.sh`** — Validates configuration files
- **`scripts/src/process.py`** — Processes input data and produces a summary report
## Workflow
1. Validate: `bash scripts/.scripts/validate.sh "$INPUT_FILE"`
2. Process: `uv run scripts/src/process.py --input results.json`
```
The same convention applies in `references/*.md` files — paths are relative to the skill root.
---
## Designing scripts for agentic use
### Hard requirement: no interactive prompts
Agents operate in non-interactive shells. A script that blocks on interactive input will hang indefinitely.
```bash
# Bad: hangs waiting for input
$ python scripts/deploy.py
Target environment: _
# Good: clear error with guidance
$ python scripts/deploy.py
Error: --env is required. Options: development, staging, production.
Usage: python scripts/deploy.py --env staging --tag v1.2.3
```
Accept all input via:
- Command-line flags (`--env staging`)
- Environment variables (`TARGET_ENV=staging`)
- Stdin (pipe-safe, non-blocking)
### Expose --help
`--help` output is the primary way an agent learns your script's interface:
```
Usage: scripts/process.py [OPTIONS] INPUT_FILE
Process input data and produce a summary report.
Options:
--format FORMAT Output format: json, csv, table (default: json)
--output FILE Write output to FILE instead of stdout
--verbose Print progress to stderr
Examples:
scripts/process.py data.csv
scripts/process.py --format csv --output report.csv data.csv
```
Keep it concise — the output enters the agent's context window.
### Write helpful error messages
```
Error: --format must be one of: json, csv, table.
Received: "xml"
```
Not: `Error: invalid input`
An opaque error wastes a turn. The message should say what went wrong, what was expected, what to try.
### Use structured output
Prefer JSON, CSV, TSV over free-form text. Structured formats can be consumed by both the agent and standard tools (`jq`, `cut`, `awk`).
```
# Hard to parse
NAME STATUS CREATED
my-service running 2025-01-15
# Machine-readable
{"name": "my-service", "status": "running", "created": "2025-01-15"}
```
Separate data from diagnostics:
- **stdout** → structured data output
- **stderr** → progress messages, warnings, diagnostics
### Further design requirements
| Requirement | Why |
|-------------|-----|
| **Idempotency** | Agents may retry commands. "Create if not exists" is safer than "create and fail on duplicate." |
| **Input validation** | Reject ambiguous input with a clear error rather than guessing. Use enums and closed sets. |
| **`--dry-run` support** | For destructive/stateful operations, let the agent preview what will happen. |
| **Meaningful exit codes** | Use distinct codes for different failure types (not found, invalid args, auth failure). Document them in `--help`. |
| **Safe defaults** | Destructive operations should require explicit flags (`--confirm`, `--force`). |
| **Predictable output size** | Agent harnesses often truncate tool output beyond ~10–30K characters. Default to summaries; support `--offset` for pagination or require `--output FILE` to opt in to large stdout. |
---
## When to bundle a script
Signal: the agent independently writes the same helper logic (a chart builder, a data parser, a validator) across multiple test runs.
When you see that pattern:
1. Extract the repeated logic into a tested script
2. Place it in `scripts/`
3. Document it in `SKILL.md` under "Available scripts"
4. Reference it with a specific run instruction
This is more reliable than letting the agent reinvent the logic each time, and it gives you a stable artifact to test and maintain.
+213
View File
@@ -0,0 +1,213 @@
# Agent Skills — Full Specification Reference
Source: https://agentskills.io/specification
## Directory structure
```
skill-name/
├── SKILL.md # Required: metadata + instructions
├── scripts/ # Optional: executable code
├── references/ # Optional: documentation loaded on demand
├── assets/ # Optional: templates, resources
└── ... # Any additional files or directories
```
## SKILL.md format
The file must contain YAML frontmatter followed by Markdown content.
---
## Frontmatter fields
| Field | Required | Constraints |
|-------|----------|-------------|
| `name` | Yes | Max 64 chars. Lowercase letters, numbers, hyphens only. No leading/trailing/consecutive hyphens. Must match the parent directory name. |
| `description` | Yes | Max 1024 chars. Non-empty. Describes what the skill does and when to use it. |
| `license` | No | License name or reference to a bundled license file. |
| `compatibility` | No | Max 500 chars. Indicates environment requirements (product, packages, network, etc.). |
| `metadata` | No | Arbitrary key-value mapping (string → string) for additional metadata. |
| `allowed-tools` | No | Space-separated string of pre-approved tools. (Experimental) |
---
## `name` field
Rules:
- 1–64 characters
- Only: lowercase letters `a-z`, digits `0-9`, hyphens `-`
- Must not start or end with a hyphen
- Must not contain consecutive hyphens `--`
- **Must match the parent directory name exactly**
Valid examples:
```yaml
name: pdf-processing
name: data-analysis
name: code-review
```
Invalid examples:
```yaml
name: PDF-Processing # uppercase not allowed
name: -pdf # cannot start with hyphen
name: pdf--processing # consecutive hyphens not allowed
```
---
## `description` field
Rules:
- 1–1024 characters
- Should describe both **what** the skill does and **when** to use it
- Include specific keywords that help agents identify relevant tasks
**Good example:**
```yaml
description: >
Extracts text and tables from PDF files, fills PDF forms, and merges
multiple PDFs. Use when working with PDF documents or when the user
mentions PDFs, forms, or document extraction.
```
**Poor example:**
```yaml
description: Helps with PDFs.
```
Principles for effective descriptions:
- Use imperative phrasing: "Use this skill when…" not "This skill does…"
- Focus on user intent, not internal mechanics
- Be explicit about indirect triggers: "even if they don't explicitly mention 'CSV'"
- Err on the side of being specific and slightly pushy
- Keep it concise — a few sentences to a short paragraph
---
## `license` field
- Specifies the license applied to the skill
- Keep it short — either the name of a license or the name of a bundled license file
Example:
```yaml
license: Apache-2.0
license: Proprietary. LICENSE.txt has complete terms.
```
---
## `compatibility` field
- 1–500 characters if provided
- Only include if your skill has specific environment requirements
- Can indicate: intended product, required system packages, network access
Examples:
```yaml
compatibility: Designed for Claude Code (or similar products)
compatibility: Requires git, docker, jq, and access to the internet
compatibility: Requires Python 3.14+ and uv
```
Most skills do not need this field.
---
## `metadata` field
- A map from string keys to string values
- Use for storing additional properties not defined by the spec
- Make key names reasonably unique to avoid conflicts
Example:
```yaml
metadata:
author: example-org
version: "1.0"
```
---
## `allowed-tools` field
- A space-separated string of tools pre-approved to run
- Experimental — support varies between agent implementations
Example:
```yaml
allowed-tools: Bash(git:*) Bash(jq:*) Read
```
---
## Body content
The Markdown body after the frontmatter contains the skill instructions.
Recommended sections:
- Step-by-step instructions
- Examples of inputs and outputs
- Common edge cases
The agent loads the entire body when the skill activates. Keep it under 500 lines. Move longer reference material to separate files.
---
## Progressive disclosure
Skills are loaded in three stages:
| Stage | Content | Size |
|-------|---------|------|
| Startup | `name` + `description` only | ~100 tokens |
| Activation | Full `SKILL.md` body | < 5,000 tokens recommended |
| On demand | Files in `scripts/`, `references/`, `assets/` | As needed |
Keep `SKILL.md` under 500 lines. Move detailed reference material to separate files and tell the agent *when* to load them — not just "see references/ for details."
---
## File references
Use relative paths from the skill root:
```markdown
See [the reference guide](references/REFERENCE.md) for details.
Run the extraction script:
scripts/extract.py
```
Keep file references one level deep from `SKILL.md`. Avoid deeply nested reference chains.
---
## Optional directories
### `scripts/`
Contains executable code agents can run. Scripts should:
- Be self-contained or clearly document dependencies
- Include helpful error messages
- Handle edge cases gracefully
Supported languages depend on the agent implementation. Common: Python, Bash, JavaScript.
### `references/`
Contains documentation agents read on demand:
- `REFERENCE.md` — Detailed technical reference
- `FORMS.md` — Form templates or structured data formats
- Domain-specific files (`finance.md`, `legal.md`, etc.)
Keep individual reference files focused. Agents load these on demand — smaller files = less context used.
### `assets/`
Contains static resources:
- Templates (document, configuration)
- Images (diagrams, examples)
- Data files (lookup tables, schemas)