2.5 KiB
2.5 KiB
Cross-client system identity and registry method
Apply this after canonical client dossiers contain structured system rows.
Observation first
Treat each client/system row as an observation, not automatically as a globally unique system. Preserve:
- exact observed label;
- client ID and exact client name;
- relationship and plain-language meaning;
- contractor/company with bounded role;
- direct evidence and confidence;
- canonical dossier path.
A generated registry must be traceable back to every observation and must not become a second editable evidence store.
Identity decisions
Use conservative tiers:
- Exact normalized named identity — case/whitespace/dash/Markdown differences may aggregate when the label is clearly a named product, domain, acronym, or shared service.
- Reviewed alias identity — merge variant labels only when an official source, unmistakable product identity, or explicit maintained rule proves equivalence.
- Client-scoped generic identity — identical phrases such as official website, institution portal, virtual exhibition, or public-service site remain separate per client.
- Composite/ambiguous identity — labels joining multiple systems remain separate until research can split them without losing the stated client relationship.
Normalization is a matching aid, not evidence. Never use fuzzy similarity alone to merge records. Preserve observed labels as aliases and explain the registry classification.
Stable generated dossiers
- Prefer hash-derived IDs/slugs from the durable identity key rather than alphabetic sequence numbers.
- Include observed aliases, distinct client count, observation count, confidence distribution, and a client-relationship table.
- Carry direct evidence through from canonical rows.
- For legacy rows without separate meaning/contractor columns, label those fields as not separately structured rather than inventing values.
- Emit machine-readable JSON and validate unique IDs/slugs, observation totals, client references, and one route per unique record.
Continuous research loop
During each recurring research batch:
- Improve canonical client rows with exact official system names, purpose, relationship, contractor role, date, evidence, and confidence.
- Review new labels against the alias rules.
- Record a merge only when equivalence is defensible; otherwise preserve separation.
- Regenerate and compare unique-system and observation counts.
- Investigate high-frequency ambiguous labels and missing contractors as priority research targets.