# ROADMAP_MATCHING vNext

Status: Draft  
Scope: Vocabulary reconciliation in the digiCULT/xTree context  
Related documents: M001, M002, M003

## Goal

The matching roadmap describes the staged development of Reconcilix vocabulary reconciliation.

This version separates three concerns:

1. **Processing Pipeline** – where in the workflow a problem is handled.
2. **Matching Capabilities** – what Reconcilix should be able to do.
3. **Matching Techniques** – how a capability may be implemented.

Evaluation documents validate individual capabilities against real data.

## Processing Pipeline

The pipeline is defined in M002 and summarized here:

| Stage | Name |
|---|---|
| P0 | Field Validation / Pre-Filtering |
| P1 | Normalization |
| P2 | Tokenization and Segmentation |
| P3 | Morphological Analysis |
| P4 | Compound Analysis |
| P5 | Semantic Interpretation |
| P6 | Candidate Generation |
| P7 | Semantic Type Filtering |
| P8 | Candidate Evaluation and Ranking |
| P9 | Explanation / Expert Review / Feedback Loop |

## Matching Capabilities

### C00 – Exact and synonym-based matching

Status: implemented  
Pipeline stages: P6, P8  
Main patterns: EXACT_MATCH

Scope:
- preferred labels
- known synonyms
- direct concept lookup

### C01 – Basic normalization

Status: partially implemented  
Pipeline stages: P1  
Main patterns: LEXICAL_VARIANT

Scope:
- case folding
- Unicode normalization
- punctuation normalization
- whitespace normalization

### C02 – Tokenization and expression segmentation

Status: partially implemented / evaluated in E001  
Pipeline stages: P2  
Main patterns: BOUND_OR_FRAGMENTED_TERM, MULTI_TYPE_EXPRESSION, COMPOSITE_TERM

Scope:
- multi-word terms
- hyphenated forms
- slash forms
- conjunction patterns
- fragmented source values

### C03 – Morphological analysis

Status: started  
Pipeline stages: P3  
Main patterns: LEXICAL_VARIANT

Scope:
- plural/singular handling
- inflection
- genitive forms
- lemmatization candidates

### C04 – Compound analysis

Status: started / evaluated in E001  
Pipeline stages: P4  
Main patterns: COMPOSITE_TERM, NAMED_ENTITY_COMPONENT

Scope:
- German compound splitting
- suffix and object-type head detection
- domain-specific compound rules

### C05 – Terminology expansion

Status: planned  
Pipeline stages: P1, P5, P6, P8  
Main patterns: LEXICAL_VARIANT, SEMANTIC_GENERALIZATION

Scope:
- synonyms
- historical spellings
- OCR variants
- curated domain variants
- controlled term expansions

### C06 – Multi-object recognition

Status: planned  
Pipeline stages: P2, P6, P8, P9  
Main patterns: MULTI_TYPE_EXPRESSION

Scope:
- recognition of multiple object types in one source value
- ensemble expressions
- multi-valued candidate output

### C07 – Named entity component analysis

Status: planned / vision  
Pipeline stages: P4, P5  
Main patterns: NAMED_ENTITY_COMPONENT, CONTEXT_DEPENDENT

Scope:
- saint names
- person names
- place names
- named artworks or named object components
- separation of proper name and object-type component

### C08 – Context-aware vocabulary reconciliation

Status: vision  
Pipeline stages: P5, P8, P9  
Main patterns: CONTEXT_DEPENDENT, NOT_AN_OBJECT_TYPE, SOURCE_DATA_ISSUE

Scope:
- TEI context
- LIDO context
- digiCULT.web structural fields
- document context
- expert feedback context

### C09 – Semantic candidate search

Status: vision  
Pipeline stages: P6  
Main patterns: SEMANTIC_GENERALIZATION, CONTEXT_DEPENDENT

Scope:
- semantic candidate retrieval
- vector search
- semantic neighbourhoods

Possible techniques:
- Qdrant
- embeddings such as BGE

### C10 – Assisted semantic interpretation

Status: vision  
Pipeline stages: P5, P8, P9  
Main patterns: CONTEXT_DEPENDENT, NAMED_ENTITY_COMPONENT, SOURCE_DATA_ISSUE, difficult SEMANTIC_GENERALIZATION cases

Scope:
- assisted interpretation
- reviewable explanation
- difficult case classification
- hybrid matching strategies

Possible techniques:
- LLMs
- rule-plus-model workflows
- expert-in-the-loop workflows

### C11 – Semantic type filtering

Status: planned  
Pipeline stages: P7, P8  
Main patterns: WRONG_SEMANTIC_CLASS, NOT_AN_OBJECT_TYPE

Scope:
- restrict candidates to intended target class
- demote incompatible candidates
- distinguish object type, title, motif, person, place, event, date, and other value types

### C12 – Expert feedback and gold-standard workflow

Status: started in E002/F01  
Pipeline stages: P9  
Main patterns: all patterns

Scope:
- expert feedback export
- matching pattern annotation
- suggested target collection
- comments
- reviewable evidence
- feedback loop into M001/M002/ROADMAP

## Mapping: previous levels to vNext capabilities

| Previous level | Previous title | vNext capability |
|---|---|---|
| Level 0 | Exact Match | C00 |
| Level 1 | Normalization | C01 |
| Level 2 | Linguistic Tokenization | C02 |
| Level 3 | Morphological Analysis | C03 |
| Level 4 | Compound Analysis | C04 |
| Level 5 | Terminology Expansion | C05 |
| Level 6 | Multi-Object Recognition | C06 |
| Level 7 | Named Entity Analysis | C07 |
| Level 8 | Context-aware Reconciliation | C08 |
| Level 9 | Semantic Candidate Search | C09 |
| Level 10 | Assisted Semantic Interpretation | C10 |
| new | Semantic Type Filtering | C11 |
| new | Expert Feedback / Gold Standard | C12 |

## Principles

- Each capability should be independently evaluable.
- Each evaluation should state which M001 pattern and which M002 pipeline stage it addresses.
- Techniques such as Qdrant, BGE, LLMs, deterministic rules, dictionaries, or OpenRefine services must be justified by pattern and pipeline stage.
- The current scope is vocabulary matching for xTree/digiCULT vocabularies.
- Entity matching against Wikidata or comparable targets is out of scope for this roadmap version, but may grow from the same methodological structure later.
