This commit is contained in:
@@ -0,0 +1,45 @@
|
||||
# Cross-client system identity and registry method
|
||||
|
||||
Apply this after canonical client dossiers contain structured system rows.
|
||||
|
||||
## Observation first
|
||||
|
||||
Treat each client/system row as an observation, not automatically as a globally unique system. Preserve:
|
||||
|
||||
- exact observed label;
|
||||
- client ID and exact client name;
|
||||
- relationship and plain-language meaning;
|
||||
- contractor/company with bounded role;
|
||||
- direct evidence and confidence;
|
||||
- canonical dossier path.
|
||||
|
||||
A generated registry must be traceable back to every observation and must not become a second editable evidence store.
|
||||
|
||||
## Identity decisions
|
||||
|
||||
Use conservative tiers:
|
||||
|
||||
1. **Exact normalized named identity** — case/whitespace/dash/Markdown differences may aggregate when the label is clearly a named product, domain, acronym, or shared service.
|
||||
2. **Reviewed alias identity** — merge variant labels only when an official source, unmistakable product identity, or explicit maintained rule proves equivalence.
|
||||
3. **Client-scoped generic identity** — identical phrases such as official website, institution portal, virtual exhibition, or public-service site remain separate per client.
|
||||
4. **Composite/ambiguous identity** — labels joining multiple systems remain separate until research can split them without losing the stated client relationship.
|
||||
|
||||
Normalization is a matching aid, not evidence. Never use fuzzy similarity alone to merge records. Preserve observed labels as aliases and explain the registry classification.
|
||||
|
||||
## Stable generated dossiers
|
||||
|
||||
- Prefer hash-derived IDs/slugs from the durable identity key rather than alphabetic sequence numbers.
|
||||
- Include observed aliases, distinct client count, observation count, confidence distribution, and a client-relationship table.
|
||||
- Carry direct evidence through from canonical rows.
|
||||
- For legacy rows without separate meaning/contractor columns, label those fields as not separately structured rather than inventing values.
|
||||
- Emit machine-readable JSON and validate unique IDs/slugs, observation totals, client references, and one route per unique record.
|
||||
|
||||
## Continuous research loop
|
||||
|
||||
During each recurring research batch:
|
||||
|
||||
1. Improve canonical client rows with exact official system names, purpose, relationship, contractor role, date, evidence, and confidence.
|
||||
2. Review new labels against the alias rules.
|
||||
3. Record a merge only when equivalence is defensible; otherwise preserve separation.
|
||||
4. Regenerate and compare unique-system and observation counts.
|
||||
5. Investigate high-frequency ambiguous labels and missing contractors as priority research targets.
|
||||
Reference in New Issue
Block a user