# Using Scripts in Skills Source: https://agentskills.io/skill-creation/using-scripts Scripts in `scripts/` let agents run executable code as part of a skill's workflow. This reference covers the folder convention, one-off commands, self-contained bundled scripts, and design principles for agentic use. --- ## Scripts folder convention ``` scripts/ ├── Taskfile.yml # Required: maps operator commands to modules ├── package.json # Required when TypeScript/JS code exists ├── tsconfig.json # Required when TypeScript code exists ├── pnpm-lock.yaml # Committed lockfile ├── .scripts/ # Shell scripts (.sh) — always placed here └── src/ # TypeScript, JavaScript, Python, or other language source ├── cli/ │ └── commands.ts # CLI dispatcher (action registry pattern) ├── actions/ │ └── / │ └── action.ts # One action per directory; exports run() └── services/ # Shared logic reused across actions ``` **Rules:** - `scripts/Taskfile.yml` is **always required** when a `scripts/` directory exists — it is the single entry point for every operator command - All `.sh` files go in `scripts/.scripts/` — never directly in `scripts/` - Shell modules follow the `Taskfile → api → lib` architecture described below - TypeScript/JS source goes in `scripts/src/` with pnpm + tsx for dev, tsc for build - Supporting config files (`package.json`, `tsconfig.json`, `pnpm-lock.yaml`, etc.) live in `scripts/` alongside the subdirectories **Run via Taskfile (preferred — all languages):** ```bash cd scripts && task validate -- --skill-dir="/path/to/skill" cd scripts && task scaffold -- --skill-name="my-skill" cd scripts && task build ``` **Run TypeScript directly with tsx (dev):** ```bash cd scripts && npx tsx src/cli/commands.ts --action validate --skill-dir /path/to/skill ``` **Run shell API directly:** ```bash bash scripts/.scripts/validator/api/skill--execute.sh --skill-dir="/path/to/skill" ``` **Run Python:** ```bash uv run scripts/src/process.py --input file.json ``` --- ## Shell module architecture (Taskfile → api → lib) When a skill ships shell scripts, organise them as **modules** rather than flat files. This is the `wrapper-first` pattern: every public command is a thin API wrapper; all logic lives in composable lib functions. ### Why this matters - **Agents run the Taskfile task**; they never need to know the internal paths - **Operators run the Taskfile task**; the api/ file is the only public surface - **Logic is testable** in isolation inside lib/; the api/ file has zero logic ### Module layout ``` scripts/ ├── Taskfile.yml # Root aggregator — includes: + TypeScript tasks └── .scripts/ └── / # e.g. validator, scaffolder, parser ├── Taskfile.yml # Module-level: defines this module's tasks ├── api/ │ └── --.sh # Thin wrapper — sources lib, calls one function └── lib/ ├── --index.sh # Composes: sources env-reader, env-validator, index-api ├── --index-api.sh # Sources every lib function file in this module ├── --env-vars-reader.sh # Reads env vars (no-op for CLI-only modules) ├── --env-vars-validator.sh # Validates required env vars (no-op for CLI-only modules) └── ---.sh # Implementation: one function per file ``` **Naming rules:** | Artefact | Convention | Example | |---|---|---| | Domain folder | `` — what the module *is* | `validator` | | API file | `--.sh` — what it *does* | `skill--execute.sh` | | Lib function file | `---.sh` | `-skill--execute.sh` | | Shell function name | `_____` | `_validator__skill__execute` | | Taskfile task | descriptive verb phrase | `validate-shell` | **Module vs action — the key distinction:** - The **module** (domain folder) is a noun describing *what the module is*: `validator`, `scaffolder`, `parser` - The **action** (api file + function suffix) is a verb describing *what it does*: `execute`, `run`, `parse`, `build` - A module named `validate` is wrong — `validate` is an action, not a module identity ### Module-level Taskfile.yml (inside .scripts//) Each shell module has its own `Taskfile.yml`. It only knows about its own actions: ```yaml # .scripts/validator/Taskfile.yml version: "3" tasks: execute: desc: -- --skill-dir="" cmds: - | ./.scripts/validator/api/skill--execute.sh {{ .CLI_ARGS }} silent: true ``` - One task per action in the module - Task names are action verbs: `execute`, `run`, `build`, `parse` - `{{ .CLI_ARGS }}` forwards all flags to the api script - `silent: true` keeps output clean ### Root Taskfile.yml (scripts/Taskfile.yml) The root Taskfile is an **aggregator** — it imports shell modules via `includes:` and adds any TypeScript/Python tasks inline: ```yaml # scripts/Taskfile.yml version: "3" includes: validator: ./.scripts/validator/Taskfile.yml # Add more modules here as they are created: # scaffolder: ./.scripts/scaffolder/Taskfile.yml tasks: default: cmds: - task --list-all silent: true validate: desc: Validate via TypeScript CLI -- --skill-dir="" cmds: - npx tsx src/cli/commands.ts --action validate {{ .CLI_ARGS }} silent: true ``` - Shell modules are namespaced automatically: `validator:execute`, `scaffolder:run` - TypeScript/Python tasks are defined inline (no sub-Taskfile needed) - `default` task runs `task --list-all` so the operator can always discover what's available - Adding a new shell module = one new line under `includes:` ### api/--.sh (thin wrapper) ```bash #!/bin/bash # Thin wrapper. No logic here. . ./.scripts/validator/lib/--index.sh _validator__skill__execute "$@" ``` - Sources `lib/--index.sh` (relative to where the task runs — `scripts/`) - Calls exactly one lib function and forwards `"$@"` - Never contains conditionals, loops, or string manipulation ### lib/--index.sh (bootstrap) ```bash #!/bin/bash . ./.scripts/validator/lib/--env-vars-reader.sh . ./.scripts/validator/lib/--env-vars-validator.sh . ./.scripts/validator/lib/--index-api.sh ``` - Fixed order: env-reader → env-validator → index-api - No logic — only `source` statements ### lib/--index-api.sh (function loader) ```bash #!/bin/bash . ./.scripts/validator/lib/-skill--execute.sh ``` - Sources every lib function file in the module - Add one line per function file; no other content ### lib/--env-vars-reader.sh and lib/--env-vars-validator.sh For CLI-only modules (all input via flags), these are no-ops: ```bash #!/bin/bash # CLI-only module — no env vars required ``` For modules that consume env vars, `--env-vars-reader.sh` exports them and `--env-vars-validator.sh` fails fast with a clear message if required vars are missing. ### lib/---.sh (implementation) ```bash #!/bin/bash _validator__skill__execute() { local skill_dir="" # ... parse flags, validate, implement } ``` - One function per file; filename mirrors the function name (minus the `_domain__` prefix) - The function contains all logic; the api wrapper has none ### Implementation workflow 1. Create the domain folder: `scripts/.scripts//api/` and `scripts/.scripts//lib/` 2. Write `lib/---.sh` with the full implementation 3. Write `lib/--index-api.sh` sourcing it 4. Write `lib/--index.sh` with the three bootstrap sources 5. Write `lib/--env-vars-reader.sh` and `lib/--env-vars-validator.sh` (even if no-op) 6. Write `api/--.sh` as the thin wrapper 7. `chmod +x` all `.sh` files in the module 8. Add the Taskfile task 9. Test: `cd scripts && task -- --flag=value` --- ## One-off commands (no scripts/ directory needed) When an existing package already does what you need, reference it directly in `SKILL.md`: | Runner | Command | Notes | |--------|---------|-------| | `uvx` | `uvx ruff@0.8.0 check .` | Python. Ships with uv. Fast, aggressive caching. | | `pipx` | `pipx run 'black==24.10.0' .` | Python. Available via OS package managers. | | `npx` | `npx eslint@9 --fix .` | Node.js packages. Ships with npm. | | `bunx` | `bunx eslint@9 --fix .` | Bun's npx equivalent. Bun-only environments. | | `go run` | `go run golang.org/x/tools/cmd/goimports@v0.28.0 .` | Go. Built into go toolchain. | **Tips:** - Pin versions (`npx eslint@9.0.0`) for reproducibility - State prerequisites in `SKILL.md` (e.g., "Requires Node.js 18+") - Move complex multi-flag commands into scripts — a tested script is more reliable than a growing one-liner --- ## Self-contained scripts with inline dependencies Bundle scripts in `scripts/`. Each script declares its own dependencies — no separate manifest or install step required. ### Python (PEP 723) — recommended ```python # scripts/process.py # /// script # dependencies = [ # "requests>=2.31", # "beautifulsoup4>=4.12,<5", # ] # requires-python = ">=3.11" # /// from bs4 import BeautifulSoup import sys # ... script content ``` Run with: ```bash uv run scripts/process.py --input data.json pipx run scripts/process.py --input data.json # alternative ``` `uv run` creates an isolated environment, installs dependencies, and runs the script. Use `uv lock --script` for a full lockfile. ### Bash — for simple shell operations ```bash #!/usr/bin/env bash # scripts/.scripts/validate.sh set -euo pipefail # ... script content ``` Run with: ```bash bash scripts/.scripts/validate.sh "$INPUT_FILE" ``` ### Deno TypeScript — self-contained by default ```typescript // scripts/extract.ts #!/usr/bin/env -S deno run import * as cheerio from "npm:cheerio@1.0.0"; // ... script content ``` Run with: `deno run scripts/extract.ts` --- ## Referencing scripts from SKILL.md Use relative paths from the skill directory root: ```markdown ## Available scripts - **`scripts/.scripts/validate.sh`** — Validates configuration files - **`scripts/src/process.py`** — Processes input data and produces a summary report ## Workflow 1. Validate: `bash scripts/.scripts/validate.sh "$INPUT_FILE"` 2. Process: `uv run scripts/src/process.py --input results.json` ``` The same convention applies in `references/*.md` files — paths are relative to the skill root. --- ## Designing scripts for agentic use ### Hard requirement: no interactive prompts Agents operate in non-interactive shells. A script that blocks on interactive input will hang indefinitely. ```bash # Bad: hangs waiting for input $ python scripts/deploy.py Target environment: _ # Good: clear error with guidance $ python scripts/deploy.py Error: --env is required. Options: development, staging, production. Usage: python scripts/deploy.py --env staging --tag v1.2.3 ``` Accept all input via: - Command-line flags (`--env staging`) - Environment variables (`TARGET_ENV=staging`) - Stdin (pipe-safe, non-blocking) ### Expose --help `--help` output is the primary way an agent learns your script's interface: ``` Usage: scripts/process.py [OPTIONS] INPUT_FILE Process input data and produce a summary report. Options: --format FORMAT Output format: json, csv, table (default: json) --output FILE Write output to FILE instead of stdout --verbose Print progress to stderr Examples: scripts/process.py data.csv scripts/process.py --format csv --output report.csv data.csv ``` Keep it concise — the output enters the agent's context window. ### Write helpful error messages ``` Error: --format must be one of: json, csv, table. Received: "xml" ``` Not: `Error: invalid input` An opaque error wastes a turn. The message should say what went wrong, what was expected, what to try. ### Use structured output Prefer JSON, CSV, TSV over free-form text. Structured formats can be consumed by both the agent and standard tools (`jq`, `cut`, `awk`). ``` # Hard to parse NAME STATUS CREATED my-service running 2025-01-15 # Machine-readable {"name": "my-service", "status": "running", "created": "2025-01-15"} ``` Separate data from diagnostics: - **stdout** → structured data output - **stderr** → progress messages, warnings, diagnostics ### Further design requirements | Requirement | Why | |-------------|-----| | **Idempotency** | Agents may retry commands. "Create if not exists" is safer than "create and fail on duplicate." | | **Input validation** | Reject ambiguous input with a clear error rather than guessing. Use enums and closed sets. | | **`--dry-run` support** | For destructive/stateful operations, let the agent preview what will happen. | | **Meaningful exit codes** | Use distinct codes for different failure types (not found, invalid args, auth failure). Document them in `--help`. | | **Safe defaults** | Destructive operations should require explicit flags (`--confirm`, `--force`). | | **Predictable output size** | Agent harnesses often truncate tool output beyond ~10–30K characters. Default to summaries; support `--offset` for pagination or require `--output FILE` to opt in to large stdout. | --- ## When to bundle a script Signal: the agent independently writes the same helper logic (a chart builder, a data parser, a validator) across multiple test runs. When you see that pattern: 1. Extract the repeated logic into a tested script 2. Place it in `scripts/` 3. Document it in `SKILL.md` under "Available scripts" 4. Reference it with a specific run instruction This is more reliable than letting the agent reinvent the logic each time, and it gives you a stable artifact to test and maintain.