Add skill-manager skill (migrated from skill-creator) with deploy action
- Skill content moved from ~/.agents/skills/skill-creator, renamed to skill-manager - New deploy CLI action: installs any skill via absolute-path symlinks (or copies) into $HOME/.agents/skills, $HOME/.claude/skills and $HOME/.cline/skills - Taskfile + shell module wrappers (task deploy / cli:deploy) - SKILL.md: deploy docs, origin-repository/origin-path metadata Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,238 @@
|
||||
# Skill Body — Best Practices
|
||||
|
||||
Source: https://agentskills.io/skill-creation/best-practices
|
||||
|
||||
## Core principle: start from real expertise
|
||||
|
||||
Ask an LLM to generate a skill without domain context → vague, generic output.
|
||||
Feed it real runbooks, API specs, code review comments, incident reports → specific, valuable skill.
|
||||
|
||||
Good source material:
|
||||
- Internal documentation, runbooks, style guides
|
||||
- API specifications, schemas, configuration files
|
||||
- Code review comments and issue trackers
|
||||
- Version control history — patches and fixes reveal real patterns
|
||||
- Real-world failure cases and their resolutions
|
||||
|
||||
---
|
||||
|
||||
## Content principles
|
||||
|
||||
### Add what the agent lacks — omit what it knows
|
||||
|
||||
Focus on what the agent *wouldn't* know without your skill:
|
||||
- Project-specific conventions
|
||||
- Domain-specific procedures
|
||||
- Non-obvious edge cases
|
||||
- The specific tools or APIs to use
|
||||
|
||||
**Too verbose:**
|
||||
```markdown
|
||||
## Extract PDF text
|
||||
PDF (Portable Document Format) files are a common file format that contains
|
||||
text, images, and other content. To extract text from a PDF, you'll need to
|
||||
use a library. pdfplumber is recommended because it handles most cases well.
|
||||
```
|
||||
|
||||
**Better:**
|
||||
```markdown
|
||||
## Extract PDF text
|
||||
Use pdfplumber. For scanned documents, fall back to pdf2image + pytesseract.
|
||||
```
|
||||
|
||||
Ask: "Would the agent get this wrong without this instruction?" If no → cut it.
|
||||
|
||||
### Provide defaults, not menus
|
||||
|
||||
When multiple tools could work, pick one and mention alternatives briefly.
|
||||
|
||||
```markdown
|
||||
<!-- Too many options — the agent will hesitate -->
|
||||
You can use pypdf, pdfplumber, PyMuPDF, or pdf2image...
|
||||
|
||||
<!-- Clear default with escape hatch -->
|
||||
Use pdfplumber:
|
||||
import pdfplumber
|
||||
|
||||
For scanned PDFs requiring OCR, use pdf2image + pytesseract instead.
|
||||
```
|
||||
|
||||
### Favor procedures over declarations
|
||||
|
||||
Teach the agent *how to approach* a class of problems, not what to produce for one specific instance.
|
||||
|
||||
```markdown
|
||||
<!-- Specific answer — only useful for this exact task -->
|
||||
Join the `orders` table to `customers` on `customer_id`, filter where
|
||||
`region = 'EMEA'`, and sum the `amount` column.
|
||||
|
||||
<!-- Reusable method — works for any analytical query -->
|
||||
1. Read the schema from `references/schema.yaml` to find relevant tables
|
||||
2. Join tables using the `_id` foreign key convention
|
||||
3. Apply filters from the user's request as WHERE clauses
|
||||
4. Aggregate numeric columns and format as a markdown table
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Effective patterns
|
||||
|
||||
### Gotchas section
|
||||
|
||||
The highest-value content in many skills. Environment-specific facts that defy reasonable assumptions — concrete corrections to mistakes the agent will make without being told.
|
||||
|
||||
```markdown
|
||||
## Gotchas
|
||||
- The `users` table uses soft deletes. Always include `WHERE deleted_at IS NULL`
|
||||
or results will include deactivated accounts.
|
||||
- The user ID is `user_id` in the database, `uid` in the auth service,
|
||||
and `accountId` in the billing API. All three refer to the same value.
|
||||
- The `/health` endpoint returns 200 even if the database connection is down.
|
||||
Use `/ready` to check full service health.
|
||||
```
|
||||
|
||||
Keep gotchas in `SKILL.md` — not in a reference file. The agent must read them before encountering the situation.
|
||||
|
||||
**When to add a gotcha:** whenever an agent makes a mistake you have to correct, add the correction here.
|
||||
|
||||
### Output format template
|
||||
|
||||
When the agent must produce a specific format, provide a template. More reliable than describing the format in prose — agents pattern-match well against concrete structures.
|
||||
|
||||
Short templates → inline in `SKILL.md`. Long templates or conditional-only templates → `assets/` and reference them:
|
||||
|
||||
```markdown
|
||||
## Report structure
|
||||
Use this template (full template in `assets/report-template.md`):
|
||||
|
||||
# [Analysis Title]
|
||||
## Executive summary
|
||||
[One-paragraph overview]
|
||||
## Key findings
|
||||
- Finding 1 with supporting data
|
||||
## Recommendations
|
||||
1. Specific actionable recommendation
|
||||
```
|
||||
|
||||
### Checklist for multi-step workflows
|
||||
|
||||
An explicit checklist helps the agent track progress and avoid skipping steps.
|
||||
|
||||
```markdown
|
||||
## Progress
|
||||
- [ ] Step 1: Analyze the form (run `scripts/analyze_form.py`)
|
||||
- [ ] Step 2: Create field mapping (edit `fields.json`)
|
||||
- [ ] Step 3: Validate mapping (run `scripts/validate_fields.py`)
|
||||
- [ ] Step 4: Fill the form (run `scripts/fill_form.py`)
|
||||
- [ ] Step 5: Verify output (run `scripts/verify_output.py`)
|
||||
```
|
||||
|
||||
### Validation loop
|
||||
|
||||
Instruct the agent to validate its own work before moving on.
|
||||
|
||||
```markdown
|
||||
## Editing workflow
|
||||
1. Make your edits
|
||||
2. Run validation: `python scripts/validate.py output/`
|
||||
3. If validation fails:
|
||||
- Review the error message
|
||||
- Fix the issues
|
||||
- Run validation again
|
||||
4. Only proceed when validation passes
|
||||
```
|
||||
|
||||
### Plan-validate-execute (for batch/destructive operations)
|
||||
|
||||
Have the agent create an intermediate plan, validate it against a source of truth, then execute.
|
||||
|
||||
```markdown
|
||||
## Form filling workflow
|
||||
1. Extract form fields: `python scripts/analyze_form.py input.pdf` → `form_fields.json`
|
||||
2. Create `field_values.json` mapping each field name to its intended value
|
||||
3. Validate: `python scripts/validate_fields.py form_fields.json field_values.json`
|
||||
(checks that field names exist, types are compatible, required fields are present)
|
||||
4. If validation fails, revise `field_values.json` and re-validate
|
||||
5. Fill: `python scripts/fill_form.py input.pdf field_values.json output.pdf`
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Calibrating prescriptiveness
|
||||
|
||||
Not every step needs the same specificity. Match it to the fragility of the task.
|
||||
|
||||
**Give the agent freedom** when multiple approaches are valid and variation is acceptable:
|
||||
```markdown
|
||||
## Code review process
|
||||
1. Check all database queries for SQL injection (use parameterized queries)
|
||||
2. Verify authentication checks on every endpoint
|
||||
3. Look for race conditions in concurrent code paths
|
||||
4. Confirm error messages don't leak internal details
|
||||
```
|
||||
|
||||
**Be prescriptive** when operations are fragile or exact sequence matters:
|
||||
```markdown
|
||||
## Database migration
|
||||
Run exactly this sequence:
|
||||
\```bash
|
||||
python scripts/migrate.py --verify --backup
|
||||
\```
|
||||
Do not modify the command or add additional flags.
|
||||
```
|
||||
|
||||
Most skills have a mix — calibrate each section independently.
|
||||
|
||||
---
|
||||
|
||||
## Sizing and progressive disclosure
|
||||
|
||||
- Keep `SKILL.md` under **500 lines / 5,000 tokens** — just core instructions the agent needs on every run
|
||||
- Move detailed reference material to `references/` files
|
||||
- Tell the agent *when* to load each file: "Read `references/api-errors.md` if the API returns a non-200 status code" — not a generic "see references/ for more"
|
||||
- Design coherent units: a skill that queries a database and formats results may be one coherent unit; one that also covers DB administration is probably too broad
|
||||
|
||||
---
|
||||
|
||||
## Refining with real execution
|
||||
|
||||
1. Run the skill against real tasks
|
||||
2. Read the execution traces (not just final outputs) — wasted steps reveal vague instructions; wrong approaches reveal instructions that don't apply
|
||||
3. Add missed corrections to the gotchas section
|
||||
4. Feed failed results + current `SKILL.md` to an LLM and ask for improvements
|
||||
|
||||
When prompting for improvements:
|
||||
- Generalize from feedback — fix the underlying issue, not the specific test case
|
||||
- Keep the skill lean — fewer better instructions outperform exhaustive rules
|
||||
- Explain the why — "Do X because Y tends to cause Z" works better than "ALWAYS do X"
|
||||
- Bundle repeated work — if the agent reinvents the same helper script every run, add it to `scripts/`
|
||||
|
||||
---
|
||||
|
||||
## .gitignore conventions
|
||||
|
||||
Every skill with a `scripts/` directory **must** have a `scripts/.gitignore`. One file, scoped to where the build output lives:
|
||||
|
||||
```gitignore
|
||||
# Node / pnpm
|
||||
node_modules/
|
||||
dist/
|
||||
.pnpm-store/
|
||||
|
||||
# Environment
|
||||
.env
|
||||
.env.local
|
||||
.env.*.local
|
||||
|
||||
# Python virtual environments
|
||||
.venv/
|
||||
__pycache__/
|
||||
*.pyc
|
||||
|
||||
# OS
|
||||
.DS_Store
|
||||
```
|
||||
|
||||
Place it at `<skill-name>/scripts/.gitignore`. Git traverses the tree and picks it up regardless of where the repo root is — no need for a skill-root `.gitignore` with `scripts/node_modules/` path prefixes.
|
||||
|
||||
**Rule:** Before the first `pnpm install` or `tsc` run, create `scripts/.gitignore`. Copy `scripts/.gitignore` from this skill as the canonical template.
|
||||
@@ -0,0 +1,146 @@
|
||||
# Description Optimization — Trigger Eval Loop
|
||||
|
||||
Source: https://agentskills.io/skill-creation/optimizing-descriptions
|
||||
|
||||
The `description` field is the sole activation trigger. Agents read only `name` + `description` at startup. An under-specified description means the skill won't trigger when it should; an over-broad description means it triggers when it shouldn't.
|
||||
|
||||
---
|
||||
|
||||
## Step 1 — Design trigger eval queries
|
||||
|
||||
Create ~20 queries: **8–10 should-trigger**, **8–10 should-not-trigger**.
|
||||
|
||||
Store them in `evals/eval_queries.json` (see `assets/evals-template.json` for the full format).
|
||||
|
||||
### Should-trigger queries
|
||||
|
||||
Vary along several axes:
|
||||
- **Phrasing**: formal, casual, with typos
|
||||
- **Explicitness**: some name the domain directly; others describe the need without naming it
|
||||
- **Detail**: terse prompts alongside context-heavy ones with file paths and column names
|
||||
- **Complexity**: single-step tasks alongside multi-step workflows
|
||||
|
||||
The most useful should-trigger queries are ones where the skill *would* help but the connection isn't obvious — these are where description wording makes the difference.
|
||||
|
||||
### Should-not-trigger queries (near-misses)
|
||||
|
||||
The most valuable negatives share keywords or concepts with your skill but need something different.
|
||||
|
||||
**Weak negatives** (obviously irrelevant — tests nothing):
|
||||
- "Write a fibonacci function"
|
||||
- "What's the weather today?"
|
||||
|
||||
**Strong negatives** (near-misses — tests precision):
|
||||
- "I need to update formulas in my Excel budget spreadsheet" — shares "spreadsheet" but needs Excel editing, not CSV analysis
|
||||
- "can you write a python script that reads a csv and uploads each row to postgres" — involves CSV, but the task is database ETL, not analysis
|
||||
|
||||
### Tips for realism
|
||||
|
||||
Include in your queries:
|
||||
- File paths (`~/Downloads/report_final_v2.xlsx`)
|
||||
- Personal context ("my manager asked me to…")
|
||||
- Specific details (column names, company names, data values)
|
||||
- Casual language, abbreviations, occasional typos
|
||||
|
||||
---
|
||||
|
||||
## Step 2 — Test trigger rates
|
||||
|
||||
Run each query through the agent with the skill installed. Observe whether the agent loads the skill's `SKILL.md`.
|
||||
|
||||
A query passes if:
|
||||
- `should_trigger: true` → skill was invoked
|
||||
- `should_trigger: false` → skill was not invoked
|
||||
|
||||
### Multiple runs (nondeterminism)
|
||||
|
||||
Run each query 3 times and compute a **trigger rate** (fraction of runs where skill was invoked).
|
||||
- Should-trigger passes if trigger rate ≥ 0.5
|
||||
- Should-not-trigger passes if trigger rate < 0.5
|
||||
|
||||
Example shell script for Claude Code:
|
||||
|
||||
```bash
|
||||
#!/bin/bash
|
||||
QUERIES_FILE="${1:?Usage: $0 <queries.json>}"
|
||||
SKILL_NAME="my-skill"
|
||||
RUNS=3
|
||||
|
||||
check_triggered() {
|
||||
local query="$1"
|
||||
claude -p "$query" --output-format json 2>/dev/null \
|
||||
| jq -e --arg skill "$SKILL_NAME" \
|
||||
'any(.messages[].content[]; .type == "tool_use" and .name == "Skill" and .input.skill == $skill)' \
|
||||
> /dev/null 2>&1
|
||||
}
|
||||
|
||||
count=$(jq length "$QUERIES_FILE")
|
||||
for i in $(seq 0 $((count - 1))); do
|
||||
query=$(jq -r ".[$i].query" "$QUERIES_FILE")
|
||||
should_trigger=$(jq -r ".[$i].should_trigger" "$QUERIES_FILE")
|
||||
triggers=0
|
||||
|
||||
for run in $(seq 1 $RUNS); do
|
||||
check_triggered "$query" && triggers=$((triggers + 1))
|
||||
done
|
||||
|
||||
jq -n \
|
||||
--arg query "$query" \
|
||||
--argjson should_trigger "$should_trigger" \
|
||||
--argjson triggers "$triggers" \
|
||||
--argjson runs "$RUNS" \
|
||||
'{query: $query, should_trigger: $should_trigger, triggers: $triggers, runs: $runs, trigger_rate: ($triggers / $runs)}'
|
||||
done | jq -s '.'
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Step 3 — Train/validation split
|
||||
|
||||
Split your ~20 queries:
|
||||
- **Train set (~60%, ~12 queries)**: guide improvements
|
||||
- **Validation set (~40%, ~8 queries)**: check whether improvements generalize
|
||||
|
||||
Keep both sets proportionally mixed (should-trigger and should-not-trigger). Fix the split across iterations.
|
||||
|
||||
---
|
||||
|
||||
## Step 4 — The optimization loop
|
||||
|
||||
1. **Evaluate** on train + validation sets
|
||||
2. **Identify failures** in train set only:
|
||||
- Should-trigger failures → description too narrow → broaden scope, add more "when to use" context
|
||||
- Should-not-trigger false-positives → description too broad → add specificity, clarify what the skill does *not* do
|
||||
3. **Revise the description**:
|
||||
- Address the general category that failed queries represent — don't add specific keywords from failed queries (overfitting)
|
||||
- If stuck after several iterations, try a structurally different approach rather than incremental tweaks
|
||||
- Check that description stays under 1024 characters
|
||||
4. **Repeat** steps 1–3 until train set passes or improvement plateaus
|
||||
5. **Select the best iteration** by validation pass rate — the best may be an earlier iteration, not the last
|
||||
|
||||
Five iterations is usually enough.
|
||||
|
||||
---
|
||||
|
||||
## Step 5 — Apply the result
|
||||
|
||||
1. Update the `description` field in `SKILL.md` frontmatter
|
||||
2. Verify it is under 1024 characters
|
||||
3. Try 5–10 fresh queries (never part of optimization) as a final sanity check
|
||||
|
||||
**Before and after example:**
|
||||
|
||||
```yaml
|
||||
# Before
|
||||
description: Process CSV files.
|
||||
|
||||
# After
|
||||
description: >
|
||||
Analyze CSV and tabular data files — compute summary statistics,
|
||||
add derived columns, generate charts, and clean messy data. Use this
|
||||
skill when the user has a CSV, TSV, or Excel file and wants to
|
||||
explore, transform, or visualize the data, even if they don't
|
||||
explicitly mention "CSV" or "analysis."
|
||||
```
|
||||
|
||||
The improved description is more specific about what the skill does (stats, derived columns, charts, cleaning) and broader about when it applies (CSV, TSV, Excel; even without explicit keywords).
|
||||
@@ -0,0 +1,418 @@
|
||||
# Using Scripts in Skills
|
||||
|
||||
Source: https://agentskills.io/skill-creation/using-scripts
|
||||
|
||||
Scripts in `scripts/` let agents run executable code as part of a skill's workflow. This reference covers the folder convention, one-off commands, self-contained bundled scripts, and design principles for agentic use.
|
||||
|
||||
---
|
||||
|
||||
## Scripts folder convention
|
||||
|
||||
```
|
||||
scripts/
|
||||
├── Taskfile.yml # Required: maps operator commands to modules
|
||||
├── package.json # Required when TypeScript/JS code exists
|
||||
├── tsconfig.json # Required when TypeScript code exists
|
||||
├── pnpm-lock.yaml # Committed lockfile
|
||||
├── .scripts/ # Shell scripts (.sh) — always placed here
|
||||
└── src/ # TypeScript, JavaScript, Python, or other language source
|
||||
├── cli/
|
||||
│ └── commands.ts # CLI dispatcher (action registry pattern)
|
||||
├── actions/
|
||||
│ └── <action>/
|
||||
│ └── action.ts # One action per directory; exports run()
|
||||
└── services/ # Shared logic reused across actions
|
||||
```
|
||||
|
||||
**Rules:**
|
||||
- `scripts/Taskfile.yml` is **always required** when a `scripts/` directory exists — it is the single entry point for every operator command
|
||||
- All `.sh` files go in `scripts/.scripts/` — never directly in `scripts/`
|
||||
- Shell modules follow the `Taskfile → api → lib` architecture described below
|
||||
- TypeScript/JS source goes in `scripts/src/` with pnpm + tsx for dev, tsc for build
|
||||
- Supporting config files (`package.json`, `tsconfig.json`, `pnpm-lock.yaml`, etc.) live in `scripts/` alongside the subdirectories
|
||||
|
||||
**Run via Taskfile (preferred — all languages):**
|
||||
```bash
|
||||
cd scripts && task validate -- --skill-dir="/path/to/skill"
|
||||
cd scripts && task scaffold -- --skill-name="my-skill"
|
||||
cd scripts && task build
|
||||
```
|
||||
|
||||
**Run TypeScript directly with tsx (dev):**
|
||||
```bash
|
||||
cd scripts && npx tsx src/cli/commands.ts --action validate --skill-dir /path/to/skill
|
||||
```
|
||||
|
||||
**Run shell API directly:**
|
||||
```bash
|
||||
bash scripts/.scripts/validator/api/skill--execute.sh --skill-dir="/path/to/skill"
|
||||
```
|
||||
|
||||
**Run Python:**
|
||||
```bash
|
||||
uv run scripts/src/process.py --input file.json
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Shell module architecture (Taskfile → api → lib)
|
||||
|
||||
When a skill ships shell scripts, organise them as **modules** rather than flat files. This is the `wrapper-first` pattern: every public command is a thin API wrapper; all logic lives in composable lib functions.
|
||||
|
||||
### Why this matters
|
||||
|
||||
- **Agents run the Taskfile task**; they never need to know the internal paths
|
||||
- **Operators run the Taskfile task**; the api/ file is the only public surface
|
||||
- **Logic is testable** in isolation inside lib/; the api/ file has zero logic
|
||||
|
||||
### Module layout
|
||||
|
||||
```
|
||||
scripts/
|
||||
├── Taskfile.yml # Root aggregator — includes: + TypeScript tasks
|
||||
└── .scripts/
|
||||
└── <module-noun>/ # e.g. validator, scaffolder, parser
|
||||
├── Taskfile.yml # Module-level: defines this module's tasks
|
||||
├── api/
|
||||
│ └── <action>--<sub-action>.sh # Thin wrapper — sources lib, calls one function
|
||||
└── lib/
|
||||
├── --index.sh # Composes: sources env-reader, env-validator, index-api
|
||||
├── --index-api.sh # Sources every lib function file in this module
|
||||
├── --env-vars-reader.sh # Reads env vars (no-op for CLI-only modules)
|
||||
├── --env-vars-validator.sh # Validates required env vars (no-op for CLI-only modules)
|
||||
└── -<action>--<sub-action>.sh # Implementation: one function per file
|
||||
```
|
||||
|
||||
**Naming rules:**
|
||||
|
||||
| Artefact | Convention | Example |
|
||||
|---|---|---|
|
||||
| Domain folder | `<module-noun>` — what the module *is* | `validator` |
|
||||
| API file | `<action>--<sub-action>.sh` — what it *does* | `skill--execute.sh` |
|
||||
| Lib function file | `-<action>--<sub-action>.sh` | `-skill--execute.sh` |
|
||||
| Shell function name | `_<module>__<action>__<sub_action>` | `_validator__skill__execute` |
|
||||
| Taskfile task | descriptive verb phrase | `validate-shell` |
|
||||
|
||||
**Module vs action — the key distinction:**
|
||||
- The **module** (domain folder) is a noun describing *what the module is*: `validator`, `scaffolder`, `parser`
|
||||
- The **action** (api file + function suffix) is a verb describing *what it does*: `execute`, `run`, `parse`, `build`
|
||||
- A module named `validate` is wrong — `validate` is an action, not a module identity
|
||||
|
||||
### Module-level Taskfile.yml (inside .scripts/<module>/)
|
||||
|
||||
Each shell module has its own `Taskfile.yml`. It only knows about its own actions:
|
||||
|
||||
```yaml
|
||||
# .scripts/validator/Taskfile.yml
|
||||
version: "3"
|
||||
|
||||
tasks:
|
||||
execute:
|
||||
desc: -- --skill-dir="<path>"
|
||||
cmds:
|
||||
- |
|
||||
./.scripts/validator/api/skill--execute.sh {{ .CLI_ARGS }}
|
||||
silent: true
|
||||
```
|
||||
|
||||
- One task per action in the module
|
||||
- Task names are action verbs: `execute`, `run`, `build`, `parse`
|
||||
- `{{ .CLI_ARGS }}` forwards all flags to the api script
|
||||
- `silent: true` keeps output clean
|
||||
|
||||
### Root Taskfile.yml (scripts/Taskfile.yml)
|
||||
|
||||
The root Taskfile is an **aggregator** — it imports shell modules via `includes:` and adds any TypeScript/Python tasks inline:
|
||||
|
||||
```yaml
|
||||
# scripts/Taskfile.yml
|
||||
version: "3"
|
||||
|
||||
includes:
|
||||
validator: ./.scripts/validator/Taskfile.yml
|
||||
# Add more modules here as they are created:
|
||||
# scaffolder: ./.scripts/scaffolder/Taskfile.yml
|
||||
|
||||
tasks:
|
||||
default:
|
||||
cmds:
|
||||
- task --list-all
|
||||
silent: true
|
||||
|
||||
validate:
|
||||
desc: Validate via TypeScript CLI -- --skill-dir="<path>"
|
||||
cmds:
|
||||
- npx tsx src/cli/commands.ts --action validate {{ .CLI_ARGS }}
|
||||
silent: true
|
||||
```
|
||||
|
||||
- Shell modules are namespaced automatically: `validator:execute`, `scaffolder:run`
|
||||
- TypeScript/Python tasks are defined inline (no sub-Taskfile needed)
|
||||
- `default` task runs `task --list-all` so the operator can always discover what's available
|
||||
- Adding a new shell module = one new line under `includes:`
|
||||
|
||||
### api/<action>--<sub-action>.sh (thin wrapper)
|
||||
|
||||
```bash
|
||||
#!/bin/bash
|
||||
# Thin wrapper. No logic here.
|
||||
. ./.scripts/validator/lib/--index.sh
|
||||
_validator__skill__execute "$@"
|
||||
```
|
||||
|
||||
- Sources `lib/--index.sh` (relative to where the task runs — `scripts/`)
|
||||
- Calls exactly one lib function and forwards `"$@"`
|
||||
- Never contains conditionals, loops, or string manipulation
|
||||
|
||||
### lib/--index.sh (bootstrap)
|
||||
|
||||
```bash
|
||||
#!/bin/bash
|
||||
. ./.scripts/validator/lib/--env-vars-reader.sh
|
||||
. ./.scripts/validator/lib/--env-vars-validator.sh
|
||||
. ./.scripts/validator/lib/--index-api.sh
|
||||
```
|
||||
|
||||
- Fixed order: env-reader → env-validator → index-api
|
||||
- No logic — only `source` statements
|
||||
|
||||
### lib/--index-api.sh (function loader)
|
||||
|
||||
```bash
|
||||
#!/bin/bash
|
||||
. ./.scripts/validator/lib/-skill--execute.sh
|
||||
```
|
||||
|
||||
- Sources every lib function file in the module
|
||||
- Add one line per function file; no other content
|
||||
|
||||
### lib/--env-vars-reader.sh and lib/--env-vars-validator.sh
|
||||
|
||||
For CLI-only modules (all input via flags), these are no-ops:
|
||||
|
||||
```bash
|
||||
#!/bin/bash
|
||||
# CLI-only module — no env vars required
|
||||
```
|
||||
|
||||
For modules that consume env vars, `--env-vars-reader.sh` exports them and `--env-vars-validator.sh` fails fast with a clear message if required vars are missing.
|
||||
|
||||
### lib/-<action>--<sub-action>.sh (implementation)
|
||||
|
||||
```bash
|
||||
#!/bin/bash
|
||||
_validator__skill__execute() {
|
||||
local skill_dir=""
|
||||
# ... parse flags, validate, implement
|
||||
}
|
||||
```
|
||||
|
||||
- One function per file; filename mirrors the function name (minus the `_domain__` prefix)
|
||||
- The function contains all logic; the api wrapper has none
|
||||
|
||||
### Implementation workflow
|
||||
|
||||
1. Create the domain folder: `scripts/.scripts/<domain>/api/` and `scripts/.scripts/<domain>/lib/`
|
||||
2. Write `lib/-<action>--<sub-action>.sh` with the full implementation
|
||||
3. Write `lib/--index-api.sh` sourcing it
|
||||
4. Write `lib/--index.sh` with the three bootstrap sources
|
||||
5. Write `lib/--env-vars-reader.sh` and `lib/--env-vars-validator.sh` (even if no-op)
|
||||
6. Write `api/<action>--<sub-action>.sh` as the thin wrapper
|
||||
7. `chmod +x` all `.sh` files in the module
|
||||
8. Add the Taskfile task
|
||||
9. Test: `cd scripts && task <action> -- --flag=value`
|
||||
|
||||
---
|
||||
|
||||
## One-off commands (no scripts/ directory needed)
|
||||
|
||||
When an existing package already does what you need, reference it directly in `SKILL.md`:
|
||||
|
||||
| Runner | Command | Notes |
|
||||
|--------|---------|-------|
|
||||
| `uvx` | `uvx ruff@0.8.0 check .` | Python. Ships with uv. Fast, aggressive caching. |
|
||||
| `pipx` | `pipx run 'black==24.10.0' .` | Python. Available via OS package managers. |
|
||||
| `npx` | `npx eslint@9 --fix .` | Node.js packages. Ships with npm. |
|
||||
| `bunx` | `bunx eslint@9 --fix .` | Bun's npx equivalent. Bun-only environments. |
|
||||
| `go run` | `go run golang.org/x/tools/cmd/goimports@v0.28.0 .` | Go. Built into go toolchain. |
|
||||
|
||||
**Tips:**
|
||||
- Pin versions (`npx eslint@9.0.0`) for reproducibility
|
||||
- State prerequisites in `SKILL.md` (e.g., "Requires Node.js 18+")
|
||||
- Move complex multi-flag commands into scripts — a tested script is more reliable than a growing one-liner
|
||||
|
||||
---
|
||||
|
||||
## Self-contained scripts with inline dependencies
|
||||
|
||||
Bundle scripts in `scripts/`. Each script declares its own dependencies — no separate manifest or install step required.
|
||||
|
||||
### Python (PEP 723) — recommended
|
||||
|
||||
```python
|
||||
# scripts/process.py
|
||||
# /// script
|
||||
# dependencies = [
|
||||
# "requests>=2.31",
|
||||
# "beautifulsoup4>=4.12,<5",
|
||||
# ]
|
||||
# requires-python = ">=3.11"
|
||||
# ///
|
||||
|
||||
from bs4 import BeautifulSoup
|
||||
import sys
|
||||
|
||||
# ... script content
|
||||
```
|
||||
|
||||
Run with:
|
||||
```bash
|
||||
uv run scripts/process.py --input data.json
|
||||
pipx run scripts/process.py --input data.json # alternative
|
||||
```
|
||||
|
||||
`uv run` creates an isolated environment, installs dependencies, and runs the script. Use `uv lock --script` for a full lockfile.
|
||||
|
||||
### Bash — for simple shell operations
|
||||
|
||||
```bash
|
||||
#!/usr/bin/env bash
|
||||
# scripts/.scripts/validate.sh
|
||||
set -euo pipefail
|
||||
|
||||
# ... script content
|
||||
```
|
||||
|
||||
Run with:
|
||||
```bash
|
||||
bash scripts/.scripts/validate.sh "$INPUT_FILE"
|
||||
```
|
||||
|
||||
### Deno TypeScript — self-contained by default
|
||||
|
||||
```typescript
|
||||
// scripts/extract.ts
|
||||
#!/usr/bin/env -S deno run
|
||||
|
||||
import * as cheerio from "npm:cheerio@1.0.0";
|
||||
|
||||
// ... script content
|
||||
```
|
||||
|
||||
Run with: `deno run scripts/extract.ts`
|
||||
|
||||
---
|
||||
|
||||
## Referencing scripts from SKILL.md
|
||||
|
||||
Use relative paths from the skill directory root:
|
||||
|
||||
```markdown
|
||||
## Available scripts
|
||||
- **`scripts/.scripts/validate.sh`** — Validates configuration files
|
||||
- **`scripts/src/process.py`** — Processes input data and produces a summary report
|
||||
|
||||
## Workflow
|
||||
1. Validate: `bash scripts/.scripts/validate.sh "$INPUT_FILE"`
|
||||
2. Process: `uv run scripts/src/process.py --input results.json`
|
||||
```
|
||||
|
||||
The same convention applies in `references/*.md` files — paths are relative to the skill root.
|
||||
|
||||
---
|
||||
|
||||
## Designing scripts for agentic use
|
||||
|
||||
### Hard requirement: no interactive prompts
|
||||
|
||||
Agents operate in non-interactive shells. A script that blocks on interactive input will hang indefinitely.
|
||||
|
||||
```bash
|
||||
# Bad: hangs waiting for input
|
||||
$ python scripts/deploy.py
|
||||
Target environment: _
|
||||
|
||||
# Good: clear error with guidance
|
||||
$ python scripts/deploy.py
|
||||
Error: --env is required. Options: development, staging, production.
|
||||
Usage: python scripts/deploy.py --env staging --tag v1.2.3
|
||||
```
|
||||
|
||||
Accept all input via:
|
||||
- Command-line flags (`--env staging`)
|
||||
- Environment variables (`TARGET_ENV=staging`)
|
||||
- Stdin (pipe-safe, non-blocking)
|
||||
|
||||
### Expose --help
|
||||
|
||||
`--help` output is the primary way an agent learns your script's interface:
|
||||
|
||||
```
|
||||
Usage: scripts/process.py [OPTIONS] INPUT_FILE
|
||||
|
||||
Process input data and produce a summary report.
|
||||
|
||||
Options:
|
||||
--format FORMAT Output format: json, csv, table (default: json)
|
||||
--output FILE Write output to FILE instead of stdout
|
||||
--verbose Print progress to stderr
|
||||
|
||||
Examples:
|
||||
scripts/process.py data.csv
|
||||
scripts/process.py --format csv --output report.csv data.csv
|
||||
```
|
||||
|
||||
Keep it concise — the output enters the agent's context window.
|
||||
|
||||
### Write helpful error messages
|
||||
|
||||
```
|
||||
Error: --format must be one of: json, csv, table.
|
||||
Received: "xml"
|
||||
```
|
||||
|
||||
Not: `Error: invalid input`
|
||||
|
||||
An opaque error wastes a turn. The message should say what went wrong, what was expected, what to try.
|
||||
|
||||
### Use structured output
|
||||
|
||||
Prefer JSON, CSV, TSV over free-form text. Structured formats can be consumed by both the agent and standard tools (`jq`, `cut`, `awk`).
|
||||
|
||||
```
|
||||
# Hard to parse
|
||||
NAME STATUS CREATED
|
||||
my-service running 2025-01-15
|
||||
|
||||
# Machine-readable
|
||||
{"name": "my-service", "status": "running", "created": "2025-01-15"}
|
||||
```
|
||||
|
||||
Separate data from diagnostics:
|
||||
- **stdout** → structured data output
|
||||
- **stderr** → progress messages, warnings, diagnostics
|
||||
|
||||
### Further design requirements
|
||||
|
||||
| Requirement | Why |
|
||||
|-------------|-----|
|
||||
| **Idempotency** | Agents may retry commands. "Create if not exists" is safer than "create and fail on duplicate." |
|
||||
| **Input validation** | Reject ambiguous input with a clear error rather than guessing. Use enums and closed sets. |
|
||||
| **`--dry-run` support** | For destructive/stateful operations, let the agent preview what will happen. |
|
||||
| **Meaningful exit codes** | Use distinct codes for different failure types (not found, invalid args, auth failure). Document them in `--help`. |
|
||||
| **Safe defaults** | Destructive operations should require explicit flags (`--confirm`, `--force`). |
|
||||
| **Predictable output size** | Agent harnesses often truncate tool output beyond ~10–30K characters. Default to summaries; support `--offset` for pagination or require `--output FILE` to opt in to large stdout. |
|
||||
|
||||
---
|
||||
|
||||
## When to bundle a script
|
||||
|
||||
Signal: the agent independently writes the same helper logic (a chart builder, a data parser, a validator) across multiple test runs.
|
||||
|
||||
When you see that pattern:
|
||||
1. Extract the repeated logic into a tested script
|
||||
2. Place it in `scripts/`
|
||||
3. Document it in `SKILL.md` under "Available scripts"
|
||||
4. Reference it with a specific run instruction
|
||||
|
||||
This is more reliable than letting the agent reinvent the logic each time, and it gives you a stable artifact to test and maintain.
|
||||
@@ -0,0 +1,213 @@
|
||||
# Agent Skills — Full Specification Reference
|
||||
|
||||
Source: https://agentskills.io/specification
|
||||
|
||||
## Directory structure
|
||||
|
||||
```
|
||||
skill-name/
|
||||
├── SKILL.md # Required: metadata + instructions
|
||||
├── scripts/ # Optional: executable code
|
||||
├── references/ # Optional: documentation loaded on demand
|
||||
├── assets/ # Optional: templates, resources
|
||||
└── ... # Any additional files or directories
|
||||
```
|
||||
|
||||
## SKILL.md format
|
||||
|
||||
The file must contain YAML frontmatter followed by Markdown content.
|
||||
|
||||
---
|
||||
|
||||
## Frontmatter fields
|
||||
|
||||
| Field | Required | Constraints |
|
||||
|-------|----------|-------------|
|
||||
| `name` | Yes | Max 64 chars. Lowercase letters, numbers, hyphens only. No leading/trailing/consecutive hyphens. Must match the parent directory name. |
|
||||
| `description` | Yes | Max 1024 chars. Non-empty. Describes what the skill does and when to use it. |
|
||||
| `license` | No | License name or reference to a bundled license file. |
|
||||
| `compatibility` | No | Max 500 chars. Indicates environment requirements (product, packages, network, etc.). |
|
||||
| `metadata` | No | Arbitrary key-value mapping (string → string) for additional metadata. |
|
||||
| `allowed-tools` | No | Space-separated string of pre-approved tools. (Experimental) |
|
||||
|
||||
---
|
||||
|
||||
## `name` field
|
||||
|
||||
Rules:
|
||||
- 1–64 characters
|
||||
- Only: lowercase letters `a-z`, digits `0-9`, hyphens `-`
|
||||
- Must not start or end with a hyphen
|
||||
- Must not contain consecutive hyphens `--`
|
||||
- **Must match the parent directory name exactly**
|
||||
|
||||
Valid examples:
|
||||
```yaml
|
||||
name: pdf-processing
|
||||
name: data-analysis
|
||||
name: code-review
|
||||
```
|
||||
|
||||
Invalid examples:
|
||||
```yaml
|
||||
name: PDF-Processing # uppercase not allowed
|
||||
name: -pdf # cannot start with hyphen
|
||||
name: pdf--processing # consecutive hyphens not allowed
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## `description` field
|
||||
|
||||
Rules:
|
||||
- 1–1024 characters
|
||||
- Should describe both **what** the skill does and **when** to use it
|
||||
- Include specific keywords that help agents identify relevant tasks
|
||||
|
||||
**Good example:**
|
||||
```yaml
|
||||
description: >
|
||||
Extracts text and tables from PDF files, fills PDF forms, and merges
|
||||
multiple PDFs. Use when working with PDF documents or when the user
|
||||
mentions PDFs, forms, or document extraction.
|
||||
```
|
||||
|
||||
**Poor example:**
|
||||
```yaml
|
||||
description: Helps with PDFs.
|
||||
```
|
||||
|
||||
Principles for effective descriptions:
|
||||
- Use imperative phrasing: "Use this skill when…" not "This skill does…"
|
||||
- Focus on user intent, not internal mechanics
|
||||
- Be explicit about indirect triggers: "even if they don't explicitly mention 'CSV'"
|
||||
- Err on the side of being specific and slightly pushy
|
||||
- Keep it concise — a few sentences to a short paragraph
|
||||
|
||||
---
|
||||
|
||||
## `license` field
|
||||
|
||||
- Specifies the license applied to the skill
|
||||
- Keep it short — either the name of a license or the name of a bundled license file
|
||||
|
||||
Example:
|
||||
```yaml
|
||||
license: Apache-2.0
|
||||
license: Proprietary. LICENSE.txt has complete terms.
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## `compatibility` field
|
||||
|
||||
- 1–500 characters if provided
|
||||
- Only include if your skill has specific environment requirements
|
||||
- Can indicate: intended product, required system packages, network access
|
||||
|
||||
Examples:
|
||||
```yaml
|
||||
compatibility: Designed for Claude Code (or similar products)
|
||||
compatibility: Requires git, docker, jq, and access to the internet
|
||||
compatibility: Requires Python 3.14+ and uv
|
||||
```
|
||||
|
||||
Most skills do not need this field.
|
||||
|
||||
---
|
||||
|
||||
## `metadata` field
|
||||
|
||||
- A map from string keys to string values
|
||||
- Use for storing additional properties not defined by the spec
|
||||
- Make key names reasonably unique to avoid conflicts
|
||||
|
||||
Example:
|
||||
```yaml
|
||||
metadata:
|
||||
author: example-org
|
||||
version: "1.0"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## `allowed-tools` field
|
||||
|
||||
- A space-separated string of tools pre-approved to run
|
||||
- Experimental — support varies between agent implementations
|
||||
|
||||
Example:
|
||||
```yaml
|
||||
allowed-tools: Bash(git:*) Bash(jq:*) Read
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Body content
|
||||
|
||||
The Markdown body after the frontmatter contains the skill instructions.
|
||||
|
||||
Recommended sections:
|
||||
- Step-by-step instructions
|
||||
- Examples of inputs and outputs
|
||||
- Common edge cases
|
||||
|
||||
The agent loads the entire body when the skill activates. Keep it under 500 lines. Move longer reference material to separate files.
|
||||
|
||||
---
|
||||
|
||||
## Progressive disclosure
|
||||
|
||||
Skills are loaded in three stages:
|
||||
|
||||
| Stage | Content | Size |
|
||||
|-------|---------|------|
|
||||
| Startup | `name` + `description` only | ~100 tokens |
|
||||
| Activation | Full `SKILL.md` body | < 5,000 tokens recommended |
|
||||
| On demand | Files in `scripts/`, `references/`, `assets/` | As needed |
|
||||
|
||||
Keep `SKILL.md` under 500 lines. Move detailed reference material to separate files and tell the agent *when* to load them — not just "see references/ for details."
|
||||
|
||||
---
|
||||
|
||||
## File references
|
||||
|
||||
Use relative paths from the skill root:
|
||||
|
||||
```markdown
|
||||
See [the reference guide](references/REFERENCE.md) for details.
|
||||
|
||||
Run the extraction script:
|
||||
scripts/extract.py
|
||||
```
|
||||
|
||||
Keep file references one level deep from `SKILL.md`. Avoid deeply nested reference chains.
|
||||
|
||||
---
|
||||
|
||||
## Optional directories
|
||||
|
||||
### `scripts/`
|
||||
|
||||
Contains executable code agents can run. Scripts should:
|
||||
- Be self-contained or clearly document dependencies
|
||||
- Include helpful error messages
|
||||
- Handle edge cases gracefully
|
||||
|
||||
Supported languages depend on the agent implementation. Common: Python, Bash, JavaScript.
|
||||
|
||||
### `references/`
|
||||
|
||||
Contains documentation agents read on demand:
|
||||
- `REFERENCE.md` — Detailed technical reference
|
||||
- `FORMS.md` — Form templates or structured data formats
|
||||
- Domain-specific files (`finance.md`, `legal.md`, etc.)
|
||||
|
||||
Keep individual reference files focused. Agents load these on demand — smaller files = less context used.
|
||||
|
||||
### `assets/`
|
||||
|
||||
Contains static resources:
|
||||
- Templates (document, configuration)
|
||||
- Images (diagrams, examples)
|
||||
- Data files (lookup tables, schemas)
|
||||
Reference in New Issue
Block a user