feat: publish vssa-clients skill
Validate skill / validate (push) Successful in 7s

This commit is contained in:
2026-08-14 12:10:47 +00:00
parent f0d4f416ce
commit fd33a2144a
9 changed files with 711 additions and 1 deletions
+14
View File
@@ -0,0 +1,14 @@
name: Validate skill
on:
push:
branches: [main]
pull_request:
jobs:
validate:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Validate skill package
run: python3 scripts/validate_skill.py
+23 -1
View File
@@ -1,3 +1,25 @@
# vssa-clients
Hermes skill for evidence-backed VSSA client, system, cloud, contractor, and procurement research.
Version-controlled Hermes skill for VSSA client, system, cloud, contractor, and procurement research.
## Source of truth
- Gitea: <https://gitea.lego-cloud.eu/vssa-v1-skills-code-agent/vssa-clients>
- Branch: `main`
- Runtime name: `vssa-clients`
- Runtime path: `/opt/data/skills/research/vssa-clients`
- Maintainer worktree: `/opt/data/vssa-clients-skill`
The runtime path is a symbolic link to this checked-out repository. This prevents a local skill copy from drifting away from Gitea.
## Update workflow
1. Fetch and fast-forward `main` before editing.
2. Modify `SKILL.md` or linked files in this repository.
3. Run `python3 scripts/validate_skill.py`.
4. Review `git diff --check` and the complete diff.
5. Commit and push to `main`.
6. Verify local/remote SHA equality and authenticated Gitea readback of `SKILL.md` plus a changed linked file.
7. Verify `skill_view(name="vssa-clients")` and the VSSA cron attachment.
Do not patch a detached runtime-only copy. Client-specific research belongs in the VSSA documentation repository; only reusable investigation procedures belong here.
+307
View File
File diff suppressed because one or more lines are too long
@@ -0,0 +1,45 @@
# Cross-client system identity and registry method
Apply this after canonical client dossiers contain structured system rows.
## Observation first
Treat each client/system row as an observation, not automatically as a globally unique system. Preserve:
- exact observed label;
- client ID and exact client name;
- relationship and plain-language meaning;
- contractor/company with bounded role;
- direct evidence and confidence;
- canonical dossier path.
A generated registry must be traceable back to every observation and must not become a second editable evidence store.
## Identity decisions
Use conservative tiers:
1. **Exact normalized named identity** — case/whitespace/dash/Markdown differences may aggregate when the label is clearly a named product, domain, acronym, or shared service.
2. **Reviewed alias identity** — merge variant labels only when an official source, unmistakable product identity, or explicit maintained rule proves equivalence.
3. **Client-scoped generic identity** — identical phrases such as official website, institution portal, virtual exhibition, or public-service site remain separate per client.
4. **Composite/ambiguous identity** — labels joining multiple systems remain separate until research can split them without losing the stated client relationship.
Normalization is a matching aid, not evidence. Never use fuzzy similarity alone to merge records. Preserve observed labels as aliases and explain the registry classification.
## Stable generated dossiers
- Prefer hash-derived IDs/slugs from the durable identity key rather than alphabetic sequence numbers.
- Include observed aliases, distinct client count, observation count, confidence distribution, and a client-relationship table.
- Carry direct evidence through from canonical rows.
- For legacy rows without separate meaning/contractor columns, label those fields as not separately structured rather than inventing values.
- Emit machine-readable JSON and validate unique IDs/slugs, observation totals, client references, and one route per unique record.
## Continuous research loop
During each recurring research batch:
1. Improve canonical client rows with exact official system names, purpose, relationship, contractor role, date, evidence, and confidence.
2. Review new labels against the alias rules.
3. Record a merge only when equivalence is defensible; otherwise preserve separation.
4. Regenerate and compare unique-system and observation counts.
5. Investigate high-frequency ambiguous labels and missing contractors as priority research targets.
+62
View File
@@ -0,0 +1,62 @@
# CVP IS organization-first procurement discovery
Use this route as an additional discovery path for Lithuanian public-procurement evidence. It is not a complete procurement history.
## Validated route
1. Open `https://viesiejipirkimai.lt/epps/viewOrganisations.do` and search the exact client legal name, then former-name and abbreviation variants where needed. The live organization form uses `POST /epps/viewOrganisations.do` with lowercase field names, including `within=template.group.ca`, `name=<exact legal name>`, and optional `shortName`, `city`, and `street`; preserve the session cookie and normal referrer when reproducing the form request. Do not invent a GET query parameter or capitalize `name` from the input element's `id="Name"`. If a byte-exact authoritative name yields no row, retry only bounded typography/legal-form variants such as removing Lithuanian quotation marks or using the portal's likely `AB`, `VšĮ`, or expanded legal-form rendering. Record which query matched.
2. Identity-check the returned legal name and available profile details. Parse complete organization-result table rows and bind each displayed legal name to the `prepareViewCAOrganisation.do?id=...` link in that same row. Do not collect every profile link from a wide HTML context and then substring-match the query: one result page or surrounding fragment can contain multiple organization/profile links, producing a false “exact” match or an ambiguous pair. Preserve the query variant, portal-displayed organization name, registration number/domain when available, and organization ID. If an exact-name query still yields multiple distinct exact rows, keep the identity unresolved until registration number, domain, or another authoritative identifier distinguishes them; never select the first result. A variant match is discovery only until reconciled with the authoritative client identity. Treat near-identical domains as a collision risk: open the homepage and corroborate its institution name/contact domain against the CVP IS profile before using any system link found there.
3. Open `https://viesiejipirkimai.lt/epps/prepareViewCAOrganisation.do?id=<organizationId>`.
4. Extract the real target of **PERŽIŪRĖTI VISUS PASKELBTUS SKELBIMUS**. It has the shape:
`https://viesiejipirkimai.lt/epps/notices/viewPublishedNotices.do?authorityId=<authorityId>&orgGroupId=<orgGroupId>`
5. Read both IDs from that link. Never assume they match or derive one arithmetically.
6. Review all result pages, preferably at the largest supported page size. The live list exposes pagination through a query key shaped `d-<table-id>-p=<page>`; derive the exact key and last-page value from the returned page's own navigation or `writePageSelection(...)` call rather than hard-coding the observed table ID. Preserve title, notice type/stage, status, language, upload date, and publication date.
7. Parse each HTML table row as a record. The notice **type** is commonly an anchor in the first cell, while the procurement **title** is plain text in the second cell—not an anchor. Therefore, anchor-text extraction alone can falsely report zero IT candidates even when a relevant title is visible. Extract and normalize all cells, keep the first-cell attachment URL, then keyword-filter the second-cell title.
8. Open promising notice-type links. They can be forced PDF attachments rather than HTML pages.
## Direct PDF behavior
A notice link commonly has this form:
`viewPublishedContractNotice.do?resourceId=<id>&documentId=<id>&noticeType=<id>&extId=<uuid>&lang=lt&noticeId=<id>&isNational=false`
Browser navigation may report `ERR_ABORTED` because the response is a download. This is not evidence of failure. Fetch the exact URL directly and verify:
- HTTP success;
- `Content-Type: application/pdf`;
- non-trivial size;
- `%PDF-` header;
- extractable text, or OCR when it is scanned.
Preserve organization ID, `authorityId`, `orgGroupId`, `resourceId`, `documentId`, `noticeId`, `noticeType`, `extId`, procedure identifier, exact title/authority, dates, status, direct URL, supporting quotation, and access date.
## Evidence semantics
Classify the PDF from its body, not from the filename or result snippet:
- market consultation/RFI → planned scope and market engagement only;
- tender/RFP/technical specification → intended procurement only;
- supplier offer → bid only;
- award notice or signed contract → selected supplier and bounded contractual role;
- amendment → only the change stated;
- implementation, production use, hosting, and acceptance/completion each require their own direct evidence.
An award that covers multiple explicitly named information systems can support one observation row per materially distinct system even when the systems share one lot and supplier. Keep the shared contract value and role clearly scoped to the joint award—do not imply that the full value belongs independently to every row. Bootstrap unknown canonical lifecycle-analysis entries for each new stable system identity before running strict registry validation; then keep hosting, original kickoff, and latest production release unknown unless separately evidenced.
For an award PDF, extract the exact buyer and system title, procedure identifier, winner, contract value, winner-selection date, contract-signature date, duration when stated, and subcontracting fields. If the body explicitly names a subcontractor together with its share or value, record that entity separately as a **subcontractor**, not as a co-winner, generic contractor, or inferred implementation partner. Preserve the award's own legal-entity spelling in the evidence record, then reconcile it conservatively with the existing contractor registry before creating a new identity. Do not treat an award's signature date or service duration as system kickoff, production release, acceptance, completion, or hosting evidence.
Map a notice to a canonical system only when it explicitly names the system/acronym or the scope is otherwise unambiguous. Similar IT terminology is insufficient.
## Confirmed examples and pitfalls
- A validated exact-name search can identify an organization even when external search is poor: organization `3350` exposed `authorityId=3350` and distinct `orgGroupId=3432`. This reinforces that both IDs must come from the profile link, not arithmetic or equality assumptions.
- An exact organization match and an opened notice list can legitimately yield no relevant IT procurement. Record the checked IDs and negative outcome for rotation continuity, but do not add a dossier row or claim procurement coverage; the organization route is incomplete.
- When rotating several organizations, preserve each exact `organizationId` and the profile-derived `(authorityId, orgGroupId)` pair in the durable rotation log—even when the notice scan yields no claim. This gives the next run a reproducible identity checkpoint without converting discovery metadata into dossier evidence.
- If an exact authoritative organization name returns no byte-exact result, state that explicitly and do not force a near-name profile. Continue bounded legal-form/former-name checks and direct system/procurement routes; absence from the organization search is not evidence that the organization has no procurement history.
- Organization/profile IDs and group IDs can differ (for example, `authorityId=1297`, `orgGroupId=1298`).
- Organization notice lists may contain hundreds of records and include unrelated purchases; scan exact system names/acronyms plus IT terms, then open candidates.
- One organization can expose separate canonical systems in adjacent notices. Do not merge them merely because the same authority procured both.
- The organization list may omit predecessor organizations, parent-authority purchases, older migrations, contracts, amendments, or other procurement stages. Continue direct procurement searches and former-name checks.
- Award and tender notices with the same title are distinct evidence objects. Open both; do not infer the award supplier from the tender.
- A directly extracted award PDF can bind a public-facing portal label to a formal backend name and bounded supplier role. Preserve both labels rather than silently renaming the client observation. For example, the title may say only “Informacinio portalo palaikymas ir vystymas” while the body explicitly scopes the work to a named information system's portal, states the contract duration, winner, value, and signature date. Use the client page to retain the public service label and the award body to state the formal system scope and maintenance/development role. This still does not prove production hosting, original development, or release dates.
- When a command-line PDF extractor is unavailable, use an isolated dependency invocation such as `uv run --with pymupdf python ...` and verify the downloaded file begins with `%PDF-`, has non-trivial size, and yields text before relying on it. This is a portable extraction fallback, not a reason to weaken document classification.
+46
View File
@@ -0,0 +1,46 @@
# Derived contractor registry method
Use this when canonical client-system observations contain contractor/provider text and the documentation needs a cross-client contractor view.
## Evidence model
Treat the contractor registry as a **derived view**, never as a second editable evidence store. Every relationship must retain:
- exact canonical client and system identity;
- explicitly named contractor/provider entity;
- bounded role as stated by the source (developer, implementer, service provider, licensor, host, support supplier, consortium member, public-sector operator, etc.);
- direct evidence URL and confidence;
- historical/current/planned context where known.
Do not turn `Not publicly identified`, “suppliers not named”, descriptive prose, or an ambiguous composite phrase into entities. Do not silently upgrade one role into another: development does not prove hosting or current support.
## Identity and stable routes
Normalize only safe textual differences for matching. Preserve the displayed source name and keep ambiguous composites separate until evidence supports a split or merge. Generate stable `CTR-*` IDs and slugs from a durable normalized identity key rather than sequence position. Sequence numbers may order generated files but must not define identity.
Before adding or changing a contractor label in a canonical observation, inspect the current derived registry and all existing near-name observations. Reuse the established exact label only when it denotes the same evidenced entity. Treat a standalone company and a composite such as `„Company“ su konsorciumo partneriais` as separate identities unless reviewed evidence supports a merge. Legal-designator changes (`UAB`, `AB`, quotation style) can create a new deterministic identity; do not introduce them casually. After regeneration, compare the entity count and inspect near-name groups. An unexpected increase is an identity-collision warning, not evidence that a new contractor was discovered.
A contractor dossier should list exact system/client relationships, bounded roles, confidence, and direct sources. A system dossier should expose its evidenced contractors, but the underlying client observation remains canonical.
## Implementation sequence
1. Write focused failing tests for named-entity extraction, unknown-placeholder exclusion, multi-contractor splitting, stable IDs/slugs, relationship aggregation, and one representative evidence-backed system.
2. Implement extraction and aggregation without editing generated pages directly.
3. Generate contractor overview/registry, individual dossiers, and machine-readable JSON.
4. Add a separate documentation collection with first-level **Overview** and expanded **Registry**, plus navbar/footer/search integration.
5. Add build invariants: JSON exists, IDs/slugs are unique, route count matches entity count, every relationship resolves to valid client/system IDs, and representative evidence is present.
6. Run focused tests, registry validation, type checking, full production build, and rendered browser checks for both a system page and contractor page.
7. After publishing, verify local/remote ref equality, exact CI run success, authenticated source readback, and deployed route behavior. An expected authentication redirect proves the access boundary is active, not page content; use successful CI publication and an authorized/backend content check when available before claiming deployed content.
## Lifecycle evidence interaction
A dated official announcement whose title/body explicitly says a named system started operating, launched, or went live can establish a production-release date at the source’s stated precision. This is different from generic page publication/update metadata. Quote the operational statement, cite corroborating announcements, and use the earliest date that explicitly establishes operation when official announcements differ. Do not infer delivery kickoff from go-live; kickoff needs separate evidence.
## Common pitfalls
- Parsing every capitalized phrase or semicolon clause as a company.
- Merging legal-name variants or consortium descriptions without reviewed identity evidence.
- Showing a contractor on a system page without retaining the source relationship.
- Presenting a historical creation supplier as the current maintainer.
- Claiming deployment verification solely from an unauthenticated redirect to an identity provider.
- Adding navigation without search indexing, generated-route validation, or machine-readable output.
+57
View File
@@ -0,0 +1,57 @@
# Lithuanian source-route notes
Use these routes as discovery and verification aids; they do not lower the claim/source standard in the main skill.
## Official institution sites
- On LRV-hosted and similar official sites, inspect `/sitemap.xml` to enumerate deep pages that menus and search engines miss. Prioritize paths for `asmens-duomenu-apsauga`, `informacines-sistemos`, `registrai`, `projektai`, `viesieji-pirkimai`, `nuostatai`, and `atviri-duomenys`.
- Open the discovered page and cite the canonical page URL. A sitemap entry proves that a page exists, not the claim inside it.
- Pages headed `... asmens duomenų valdytoja` or `... asmens duomenų tvarkytoja` can directly establish a controller/processor role for listed systems. Do **not** translate `valdytoja` on a personal-data page into system ownership unless system regulations or another source explicitly assign ownership.
## Official-site internal search fallback
- When external engines do not expose an older institution or article, try the official municipality/institution search route directly. A recurring Lithuanian pattern is `/search?q=<URL-encoded exact institution name>`.
- Treat the search page only as discovery. Open and cite the resulting official article; verify that the article body, not merely a search snippet, names the institution and supports the claimed purpose or relationship.
- This route is particularly useful for schools with retired domains, renamed institutions, and municipality-hosted news archives.
## data.gov.lt fallback for blocked institution sites
- When an institution's official site remains blocked after one browser-equivalent retry, search `https://data.gov.lt/datasets/?q=<URL-encoded exact institution name>`. Inspect the organization/creator facets and result cards, then open the direct organization page (`/orgs/<id>/`) and each relevant dataset record (`/datasets/<id>/`).
- The organization page can corroborate the exact publisher identity, supervising jurisdiction, and sector. A direct dataset record can establish that the institution publishes or maintains data from a named register or system when its metadata says so.
- Treat the query page and facets as discovery, not final evidence. Cite the opened organization or dataset record. Dataset publication does not by itself prove software ownership, development, maintenance, hosting, or cloud provider; preserve those as unknown unless another source assigns the role.
## CPVA search form
- CPVA's `https://cpva.lt/paieska` uses Search & Filter Pro. The live, validated result URL is `https://cpva.lt/paieska?_sf_s=<URL-encoded phrase>`. The browser rewrites the form state to this URL and the server-rendered `.search-filter-results-3173` container contains the result cards. An ordinary `POST` with `_sf_search[]` can return an apparently successful but unfiltered/empty page; do not use that as evidence that a query ran.
- Validate automation against a positive control such as the generic term `sistema` and a negative/random control. Confirm the result container changes, extract result-card titles and links, and recognize the explicit `Rezultatų nėra – įveskite / pakoreguokite paieškos frazę` empty state.
- CPVA search behaves as broad token matching rather than reliable exact-phrase matching. Long system labels can return unrelated pages sharing generic words, and even every queried system may appear to have a raw hit. Treat candidate counts as noisy discovery output; rank pages by exact acronym/name occurrence and open the source body before mapping anything.
- Search pages remain discovery evidence only. Open the resulting CPVA article, project, or attachment and require its body to prove the exact client/system relationship, project, date, funding, supplier, or delivery role before citing it.
## Official policy PDFs as SaaS evidence
- Inspect sitemap URLs for direct PDFs as well as HTML pages. School and care-institution sitemaps often expose electronic-diary usage rules, data-processing policies, or director-approved procedures that navigation menus omit.
- Extract and inspect the full PDF text. A dated official procedure can directly name the institution, product, provider, system purpose, administrative roles, and login URL; embedded PDF link annotations may also reveal a product URL not obvious in extracted prose.
- Bound the date carefully: a policy proves use or an assigned provider role at the policy date, not an unchanged current subscription, current maintenance allocation, contract term, or production hosting. Keep those fields unresolved unless a current contract or equivalent source establishes them.
- When the official policy names both the SaaS product and company, it can support the client-product relationship and provider attribution in one primary source. Open the provider's current legal/product page separately to corroborate the company's identity and product role, but do not inflate that corroboration into implementation, hosting, or contract claims.
## Lithuanian public-procurement documents
- Search results may expose direct documents at URLs shaped like `https://viesiejipirkimai.lt/epps/cft/downloadContractDocument.do?resourceId=...&documentId=...`.
- Official institution sites may also place a child `/sutartys/` page under a named system or service page. Inspect its raw anchors: these catalogues can separately link a preliminary agreement, amendment, main contract, and supplier offer even when the visible page contains almost no supplier detail. Resolve relative attachment URLs against the page URL, download each relevant body, and classify each document independently.
- Open and inspect the document body. Endpoint names and link labels such as `downloadContractDocument`, `Pagrindinė sutartis`, or `Tiekėjo pateiktas pasiūlymas` describe the expected document class but are not proof of supplier identity, signature, dates, or delivery until the body confirms them.
- If an HTML catalogue opens but an attachment rejects a generic downloader, retry once with the same browser-equivalent `User-Agent`, `Accept`, and language headers used for bot-sensitive public pages. If the body still cannot be opened, cite the catalogue only for the existence of procurement documents and leave supplier attribution unresolved.
- Classify evidence by its contents:
- technical specification or tender conditions → intended procurement and required scope;
- supplier offer → bidder and proposed scope, not award by itself;
- preliminary/framework agreement → framework parties and scope, not necessarily a call-off;
- main/signed contract → contracted role and dates;
- award notice → selected supplier, subject to identity/date alignment;
- acceptance/completion record → delivered or accepted work.
- Technical specifications are especially useful for current-system inventories, acronyms, integrations, migration scope, and planned functionality, but they do not identify the winning supplier unless the body explicitly does so.
## Pre-delivery link verification
- Extract every URL from changed dossiers and open each one before commit. Correct path truncation, stale slugs, and redirects before assigning confidence.
- Inspect redirect chains, not only the initial URL's status. An official institution page may redirect to a retired or unreachable product domain: the redirect can still prove that the institution links or historically linked the named service, but it does **not** prove that the service is currently operational. State the availability gap explicitly and avoid an unqualified current-use claim.
- For a multi-file batch, check unique evidence URLs concurrently with bounded timeouts, then retry flagged links once with a different network strategy (for example IPv4-only `curl -4`) before classifying them. Preserve genuine HTTP failures and unreachable final destinations as review items rather than silently accepting the source URL.
- Re-run the dossier table/schema check after any textual patch; a targeted replacement can accidentally remove a Markdown table cell while leaving the prose readable.
@@ -0,0 +1,135 @@
# System infrastructure and lifecycle analysis
Use this method when a cross-client registry needs per-system answers for hosting location, system kickoff, and latest production release.
## Scope boundary
Keep these facts in canonical **system-level metadata**, separate from client-observation rows:
- Client rows prove that one organization owns, operates, uses, procures, or depends on a system.
- System metadata describes the unique system as a whole.
- A client's adoption date, tenant hosting arrangement, or modernization project must not become a shared-system fact unless the source explicitly applies it to the system globally.
Generated system dossiers remain derived views and must never become editable evidence stores.
## Required fields
Every unique system record should contain all three fields, even when unknown:
1. `hosting_environment`
- known models: `cloud`, `on-premises`, `hybrid`
- record the named provider or environment when the source states it
2. `kickoff_date`
- original system delivery or implementation start
- precision: `year`, `month`, or `day`
3. `latest_release_date`
- latest directly evidenced production release or deployment
- precision: `year`, `month`, or `day`
Allowed status values are `known`, `unknown`, and `not-applicable`. Unknown and not-applicable values require a concise reason. Known values require direct evidence.
## Evidence object
For each known value, preserve at least:
```json
{
"url": "https://official.example/source",
"title": "Official source title",
"quote": "Exact text supporting the asserted value"
}
```
The quote is a claim guard: it makes reviewers check whether the source really says *production hosting*, *kickoff*, or *release*, rather than merely mentioning a nearby date or technology.
## Meaning of the fields
### Hosting environment
Accept only explicit production placement evidence. A source may identify a commercial cloud, government cloud, private cloud, institutional data centre, supplier data centre, on-premises deployment, or hybrid arrangement.
Do not infer hosting from:
- DNS, IP ownership, TLS, CDN, or public website hosting;
- eligibility for centralized IT/cloud services;
- a supplier's generic AWS/Azure/GCP capability;
- a SaaS product name without an institution/system-specific hosting statement;
- development, test, backup, or disaster-recovery infrastructure when production is not identified.
### System kickoff
Use the original system-delivery or implementation start only when the source describes it as commencement, start, kickoff, or equivalent. Do not silently substitute:
- tender publication;
- contract signature;
- funding approval;
- public launch;
- modernization start;
- project completion.
If only a later modernization kickoff is known, keep the original system kickoff unknown and preserve the modernization date in its proper project/observation context.
### Latest production release
Use an explicit production release, deployment, go-live, or version release date. A dated official announcement whose title/body explicitly says the system **started operating**, **launched**, **went live**, or was **deployed to production** may use the announcement's stated publication date as the release date: the operational statement establishes the event and the official date anchors it. Prefer the earliest authoritative announcement when corroborating reposts differ by a day, preserve all corroborating sources, and do not treat a later repost as a later release.
Do not substitute:
- a webpage updated or publication date when the body does not explicitly establish production operation;
- contract end or project completion;
- acceptance date unless the source also proves production deployment;
- latest tender or maintenance contract date.
## Identity-safe canonical storage
Prefer a canonical metadata file keyed by the registry's stable system ID and include the underlying identity key as a guard:
```json
{
"schema_version": 1,
"systems": {
"SYS-XXXXXXXXXX": {
"identity_key": "name:canonical identity",
"hosting_environment": {"status": "unknown", "reason": "...", "evidence": []},
"kickoff_date": {"status": "unknown", "reason": "...", "evidence": []},
"latest_release_date": {"status": "unknown", "reason": "...", "evidence": []}
}
}
}
```
Require exact coverage of all generated system IDs. Missing records must fail validation rather than silently defaulting to unknown, because omission and researched-but-unknown are different states.
When aliases, normalization, or client-scoped labels change, generated IDs may change. A bootstrap helper may add missing unknown skeletons, but it must:
- never overwrite researched values;
- stop on stale IDs;
- require manual evidence migration for merges and splits;
- preserve the `identity_key` check.
## Validation rules
Fail generation/build when:
- metadata IDs differ from generated registry IDs;
- an identity key does not match;
- any required field is absent;
- a status or hosting model is outside its controlled vocabulary;
- a known value lacks a direct HTTP(S) source, title, or quote;
- an unknown/not-applicable value lacks a reason;
- a date does not match its declared precision;
- generated JSON omits the analysis fields or retains an old schema version.
Publish canonical metadata as a byte-identical raw artifact with a checksum. Include the analysis in machine-readable registry JSON and generated system pages.
## Research rotation
A recurring research job should, for every unique system touched in a client batch:
1. Search explicitly for production hosting, original kickoff, and latest production release.
2. Open and classify every candidate source.
3. Update only directly supported fields.
4. Preserve explicit unknowns for the rest.
5. Regenerate, validate, build, inspect the complete diff, and verify published page and JSON output.
Initial migration should create explicit unknown records for full coverage. It must not mine existing prose automatically and turn nearby cloud or date mentions into system-level facts.
+22
View File
@@ -0,0 +1,22 @@
#!/usr/bin/env python3
from pathlib import Path
import re
ROOT = Path(__file__).resolve().parents[1]
skill = ROOT / "SKILL.md"
text = skill.read_text(encoding="utf-8")
assert text.startswith("---\n"), "frontmatter must start at byte zero"
match = re.match(r"---\n(.*?)\n---\n(.+)", text, re.S)
assert match, "invalid frontmatter/body framing"
frontmatter, body = match.groups()
assert re.search(r"^name:\s*vssa-clients\s*$", frontmatter, re.M), "wrong skill name"
description = re.search(r'^description:\s*["\']?(.*?)["\']?\s*$', frontmatter, re.M)
assert description and 0 < len(description.group(1)) <= 1024, "invalid description"
assert len(text) <= 100_000, "SKILL.md exceeds Hermes limit"
assert body.strip(), "skill body is empty"
links = re.findall(r"\[[^]]+\]\((references/[^)]+)\)", text)
assert links, "no linked references"
unique_links = sorted(set(links))
missing = [link for link in unique_links if not (ROOT / link).is_file()]
assert not missing, f"missing linked references: {missing}"
print(f"OK skill=vssa-clients chars={len(text)} references={len(unique_links)}")