- Skill content moved from ~/.agents/skills/skill-creator, renamed to skill-manager - New deploy CLI action: installs any skill via absolute-path symlinks (or copies) into $HOME/.agents/skills, $HOME/.claude/skills and $HOME/.cline/skills - Taskfile + shell module wrappers (task deploy / cli:deploy) - SKILL.md: deploy docs, origin-repository/origin-path metadata Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
5.4 KiB
Description Optimization — Trigger Eval Loop
Source: https://agentskills.io/skill-creation/optimizing-descriptions
The description field is the sole activation trigger. Agents read only name + description at startup. An under-specified description means the skill won't trigger when it should; an over-broad description means it triggers when it shouldn't.
Step 1 — Design trigger eval queries
Create ~20 queries: 8–10 should-trigger, 8–10 should-not-trigger.
Store them in evals/eval_queries.json (see assets/evals-template.json for the full format).
Should-trigger queries
Vary along several axes:
- Phrasing: formal, casual, with typos
- Explicitness: some name the domain directly; others describe the need without naming it
- Detail: terse prompts alongside context-heavy ones with file paths and column names
- Complexity: single-step tasks alongside multi-step workflows
The most useful should-trigger queries are ones where the skill would help but the connection isn't obvious — these are where description wording makes the difference.
Should-not-trigger queries (near-misses)
The most valuable negatives share keywords or concepts with your skill but need something different.
Weak negatives (obviously irrelevant — tests nothing):
- "Write a fibonacci function"
- "What's the weather today?"
Strong negatives (near-misses — tests precision):
- "I need to update formulas in my Excel budget spreadsheet" — shares "spreadsheet" but needs Excel editing, not CSV analysis
- "can you write a python script that reads a csv and uploads each row to postgres" — involves CSV, but the task is database ETL, not analysis
Tips for realism
Include in your queries:
- File paths (
~/Downloads/report_final_v2.xlsx) - Personal context ("my manager asked me to…")
- Specific details (column names, company names, data values)
- Casual language, abbreviations, occasional typos
Step 2 — Test trigger rates
Run each query through the agent with the skill installed. Observe whether the agent loads the skill's SKILL.md.
A query passes if:
should_trigger: true→ skill was invokedshould_trigger: false→ skill was not invoked
Multiple runs (nondeterminism)
Run each query 3 times and compute a trigger rate (fraction of runs where skill was invoked).
- Should-trigger passes if trigger rate ≥ 0.5
- Should-not-trigger passes if trigger rate < 0.5
Example shell script for Claude Code:
#!/bin/bash
QUERIES_FILE="${1:?Usage: $0 <queries.json>}"
SKILL_NAME="my-skill"
RUNS=3
check_triggered() {
local query="$1"
claude -p "$query" --output-format json 2>/dev/null \
| jq -e --arg skill "$SKILL_NAME" \
'any(.messages[].content[]; .type == "tool_use" and .name == "Skill" and .input.skill == $skill)' \
> /dev/null 2>&1
}
count=$(jq length "$QUERIES_FILE")
for i in $(seq 0 $((count - 1))); do
query=$(jq -r ".[$i].query" "$QUERIES_FILE")
should_trigger=$(jq -r ".[$i].should_trigger" "$QUERIES_FILE")
triggers=0
for run in $(seq 1 $RUNS); do
check_triggered "$query" && triggers=$((triggers + 1))
done
jq -n \
--arg query "$query" \
--argjson should_trigger "$should_trigger" \
--argjson triggers "$triggers" \
--argjson runs "$RUNS" \
'{query: $query, should_trigger: $should_trigger, triggers: $triggers, runs: $runs, trigger_rate: ($triggers / $runs)}'
done | jq -s '.'
Step 3 — Train/validation split
Split your ~20 queries:
- Train set (~60%, ~12 queries): guide improvements
- Validation set (~40%, ~8 queries): check whether improvements generalize
Keep both sets proportionally mixed (should-trigger and should-not-trigger). Fix the split across iterations.
Step 4 — The optimization loop
- Evaluate on train + validation sets
- Identify failures in train set only:
- Should-trigger failures → description too narrow → broaden scope, add more "when to use" context
- Should-not-trigger false-positives → description too broad → add specificity, clarify what the skill does not do
- Revise the description:
- Address the general category that failed queries represent — don't add specific keywords from failed queries (overfitting)
- If stuck after several iterations, try a structurally different approach rather than incremental tweaks
- Check that description stays under 1024 characters
- Repeat steps 1–3 until train set passes or improvement plateaus
- Select the best iteration by validation pass rate — the best may be an earlier iteration, not the last
Five iterations is usually enough.
Step 5 — Apply the result
- Update the
descriptionfield inSKILL.mdfrontmatter - Verify it is under 1024 characters
- Try 5–10 fresh queries (never part of optimization) as a final sanity check
Before and after example:
# Before
description: Process CSV files.
# After
description: >
Analyze CSV and tabular data files — compute summary statistics,
add derived columns, generate charts, and clean messy data. Use this
skill when the user has a CSV, TSV, or Excel file and wants to
explore, transform, or visualize the data, even if they don't
explicitly mention "CSV" or "analysis."
The improved description is more specific about what the skill does (stats, derived columns, charts, cleaning) and broader about when it applies (CSV, TSV, Excel; even without explicit keywords).