# Description Optimization — Trigger Eval Loop Source: https://agentskills.io/skill-creation/optimizing-descriptions The `description` field is the sole activation trigger. Agents read only `name` + `description` at startup. An under-specified description means the skill won't trigger when it should; an over-broad description means it triggers when it shouldn't. --- ## Step 1 — Design trigger eval queries Create ~20 queries: **8–10 should-trigger**, **8–10 should-not-trigger**. Store them in `evals/eval_queries.json` (see `assets/evals-template.json` for the full format). ### Should-trigger queries Vary along several axes: - **Phrasing**: formal, casual, with typos - **Explicitness**: some name the domain directly; others describe the need without naming it - **Detail**: terse prompts alongside context-heavy ones with file paths and column names - **Complexity**: single-step tasks alongside multi-step workflows The most useful should-trigger queries are ones where the skill *would* help but the connection isn't obvious — these are where description wording makes the difference. ### Should-not-trigger queries (near-misses) The most valuable negatives share keywords or concepts with your skill but need something different. **Weak negatives** (obviously irrelevant — tests nothing): - "Write a fibonacci function" - "What's the weather today?" **Strong negatives** (near-misses — tests precision): - "I need to update formulas in my Excel budget spreadsheet" — shares "spreadsheet" but needs Excel editing, not CSV analysis - "can you write a python script that reads a csv and uploads each row to postgres" — involves CSV, but the task is database ETL, not analysis ### Tips for realism Include in your queries: - File paths (`~/Downloads/report_final_v2.xlsx`) - Personal context ("my manager asked me to…") - Specific details (column names, company names, data values) - Casual language, abbreviations, occasional typos --- ## Step 2 — Test trigger rates Run each query through the agent with the skill installed. Observe whether the agent loads the skill's `SKILL.md`. A query passes if: - `should_trigger: true` → skill was invoked - `should_trigger: false` → skill was not invoked ### Multiple runs (nondeterminism) Run each query 3 times and compute a **trigger rate** (fraction of runs where skill was invoked). - Should-trigger passes if trigger rate ≥ 0.5 - Should-not-trigger passes if trigger rate < 0.5 Example shell script for Claude Code: ```bash #!/bin/bash QUERIES_FILE="${1:?Usage: $0 }" SKILL_NAME="my-skill" RUNS=3 check_triggered() { local query="$1" claude -p "$query" --output-format json 2>/dev/null \ | jq -e --arg skill "$SKILL_NAME" \ 'any(.messages[].content[]; .type == "tool_use" and .name == "Skill" and .input.skill == $skill)' \ > /dev/null 2>&1 } count=$(jq length "$QUERIES_FILE") for i in $(seq 0 $((count - 1))); do query=$(jq -r ".[$i].query" "$QUERIES_FILE") should_trigger=$(jq -r ".[$i].should_trigger" "$QUERIES_FILE") triggers=0 for run in $(seq 1 $RUNS); do check_triggered "$query" && triggers=$((triggers + 1)) done jq -n \ --arg query "$query" \ --argjson should_trigger "$should_trigger" \ --argjson triggers "$triggers" \ --argjson runs "$RUNS" \ '{query: $query, should_trigger: $should_trigger, triggers: $triggers, runs: $runs, trigger_rate: ($triggers / $runs)}' done | jq -s '.' ``` --- ## Step 3 — Train/validation split Split your ~20 queries: - **Train set (~60%, ~12 queries)**: guide improvements - **Validation set (~40%, ~8 queries)**: check whether improvements generalize Keep both sets proportionally mixed (should-trigger and should-not-trigger). Fix the split across iterations. --- ## Step 4 — The optimization loop 1. **Evaluate** on train + validation sets 2. **Identify failures** in train set only: - Should-trigger failures → description too narrow → broaden scope, add more "when to use" context - Should-not-trigger false-positives → description too broad → add specificity, clarify what the skill does *not* do 3. **Revise the description**: - Address the general category that failed queries represent — don't add specific keywords from failed queries (overfitting) - If stuck after several iterations, try a structurally different approach rather than incremental tweaks - Check that description stays under 1024 characters 4. **Repeat** steps 1–3 until train set passes or improvement plateaus 5. **Select the best iteration** by validation pass rate — the best may be an earlier iteration, not the last Five iterations is usually enough. --- ## Step 5 — Apply the result 1. Update the `description` field in `SKILL.md` frontmatter 2. Verify it is under 1024 characters 3. Try 5–10 fresh queries (never part of optimization) as a final sanity check **Before and after example:** ```yaml # Before description: Process CSV files. # After description: > Analyze CSV and tabular data files — compute summary statistics, add derived columns, generate charts, and clean messy data. Use this skill when the user has a CSV, TSV, or Excel file and wants to explore, transform, or visualize the data, even if they don't explicitly mention "CSV" or "analysis." ``` The improved description is more specific about what the skill does (stats, derived columns, charts, cleaning) and broader about when it applies (CSV, TSV, Excel; even without explicit keywords).