# ME002 — Vocabulary Matching Pipeline

Status: Draft  
Version: v0.5  
Scope: Reconcilix Processing Model für Vocabulary Matching im digiCULT/xTree-Kontext  
Related: ME001 Vocabulary Matching Patterns

## 1. Purpose

Dieses Dokument beschreibt, wie Reconcilix `SourceValue` verarbeitet, interpretiert und gegen kontrollierte Vokabulare reconciliert.

ME002 stellt den Bezug zwischen der unabhängigen Pattern-Taxonomie aus ME001 und der konkreten Verarbeitung in Reconcilix her.

Der zentrale Gedanke lautet:

> Reconcilix does not immediately search for the best CandidateConcept. Reconcilix first tries to identify an appropriate Matching Strategy for interpreting the SourceValue.

Oder kürzer:

> Reconcilix reconciles interpretations of source values against controlled vocabularies.

## 2. Core Domain Model

Reconcilix verwendet folgendes fachliches Modell:

```text
SourceValue
→ InterpretationGraph
→ InterpretationNode
→ ReconciliationResult
→ CandidateConcept
→ MatchDecision
```

### 2.1 SourceValue

Ein `SourceValue` ist der ursprüngliche Eingabewert, der reconciliert werden soll.

Eigenschaften:

- unveränderlich
- referenziert das Quellfeld und den Quelldatensatz
- kann ein einzelner Feldwert, ein Textausschnitt oder ein längerer String sein
- wird nicht überschrieben, sondern durch `InterpretationNode` weiterverarbeitet

Beispiel:

```text
SourceValue.value = "Altrarretabel"
```

### 2.2 InterpretationGraph

Ein `InterpretationGraph` enthält alle Interpretationen, die zu einem `SourceValue` erzeugt wurden.

Er ist notwendig, weil ein `SourceValue` mehrere Deutungen, Spans, Normalisierungen oder Ableitungen besitzen kann.

Beispiele:

```text
Altrarretabel
→ Altarretabel
→ Altar + Retabel
→ Retabel
```

oder:

```text
Der Herr von Bogen ging mit dem Bogen über den Bogen
→ Bogen [Span 1]
→ Bogen [Span 2]
→ Bogen [Span 3]
```

### 2.3 InterpretationNode

Ein `InterpretationNode` ist ein Knoten im `InterpretationGraph`.

Er kann darstellen:

- den unveränderten ursprünglichen Wert
- eine orthographisch normalisierte Form
- einen isolierten Textspan
- eine zerlegte Komponente
- eine semantisch interpretierte Form
- eine kontextuell angereicherte Interpretation

Ein `InterpretationNode` kann weitere `InterpretationNode` erzeugen.

### 2.4 ReconciliationResult

Ein `ReconciliationResult` beschreibt das Ergebnis eines Reconciliation-Versuchs für einen `InterpretationNode`.

Es enthält insbesondere:

- `success`
- `VocabularyMatchingPattern`
- `confidence`
- `notes`
- ggf. Hinweise auf fehlenden Kontext oder nächste Matching Strategy

Damit wird das `Vocabulary Matching Pattern` nicht als Eigenschaft des `SourceValue`, sondern als Ergebnis einer Interpretation/Reconciliation verstanden.

### 2.5 CandidateConcept

Ein `CandidateConcept` ist ein vorgeschlagener Zielbegriff aus einem kontrollierten Vokabular.

Eigenschaften:

- URI / Identifier
- preferred Term / Label
- Concept Scheme
- Score
- Herkunft des Kandidaten, z. B. exact lookup, broader search, OpenRefine, Qdrant, external Concept Scheme

### 2.6 MatchDecision

Eine `MatchDecision` ist die fachliche Entscheidung zu einem `CandidateConcept` oder zu einem `ReconciliationResult`.

Typische Entscheidungen:

- `accept`
- `reject`
- `needs_context`
- `manual_review`
- `suggest_target`

Eine `MatchDecision` kann automatisch, halbautomatisch oder durch Expert:innen erfolgen.

## 3. UML Class Diagram

```mermaid
classDiagram
    class SourceValue {
        +id
        +value
        +sourceField
        +sourceRecordId
        +datatype
        +language
        +provenance
    }

    class InterpretationGraph {
        +id
        +sourceValueId
    }

    class InterpretationNode {
        +id
        +value
        +level
        +spanStart
        +spanEnd
        +method
        +status
        +note
    }

    class ReconciliationResult {
        +id
        +success
        +pattern
        +confidence
        +note
    }

    class CandidateConcept {
        +id
        +uri
        +prefLabel
        +conceptScheme
        +score
        +source
    }

    class MatchDecision {
        +id
        +decision
        +confidence
        +decidedBy
        +comment
        +timestamp
    }

    SourceValue "1" --> "1" InterpretationGraph : has
    InterpretationGraph "1" --> "1..*" InterpretationNode : contains
    InterpretationNode "0..1" --> "0..*" InterpretationNode : derives
    InterpretationNode "1" --> "0..*" ReconciliationResult : produces
    ReconciliationResult "1" --> "0..*" CandidateConcept : contains
    CandidateConcept "0..1" --> "0..*" MatchDecision : receives
    ReconciliationResult "0..1" --> "0..*" MatchDecision : receives
```

## 4. Processing Model

Reconcilix verarbeitet einen `SourceValue` nicht linear, sondern iterativ über einen `InterpretationGraph`.

```text
SourceValue
↓
Create initial InterpretationNode
↓
Try Reconciliation
↓
Classify ReconciliationResult using Vocabulary Matching Pattern
↓
If success is sufficient:
    create MatchDecision or return CandidateConcepts
Else:
    select next Matching Strategy
    create additional InterpretationNode
    try Reconciliation again
```

## 5. Activity Diagram

```mermaid
flowchart TD
    A[SourceValue] --> B[Create InterpretationGraph]
    B --> C[Create initial InterpretationNode]
    C --> D[Try Reconciliation]
    D --> E[Create ReconciliationResult]
    E --> F{success sufficient?}

    F -->|yes| G[CandidateConcept accepted or prepared for MatchDecision]
    G --> H[MatchDecision]

    F -->|no| I[Classify Vocabulary Matching Pattern]
    I --> J[Select next Matching Strategy]
    J --> K[Create new InterpretationNode]
    K --> D

    I -->|CONTEXT_DEPENDENT| L[Enrich with context]
    L --> K
```

## 6. Matching Strategy

Eine `Matching Strategy` beschreibt, wie Reconcilix ausgehend von einem `ReconciliationResult` weiter vorgeht.

Beispiele:

| Observed Pattern | Possible Matching Strategy |
|---|---|
| `ORTHOGRAPHIC_VARIANT` | create normalized InterpretationNode |
| `LEXICAL_VARIANT` | apply singularization or lexical normalization |
| `COMPOSITE_TERM` | split or analyze compound |
| `MULTI_CONCEPT_EXPRESSION` | create multiple InterpretationNodes for components |
| `SEMANTIC_GENERALIZATION` | search broader concepts |
| `CONTEXT_DEPENDENT` | enrich with TEI, LIDO, field context, geo data, or other context |
| `NOT_AN_OBJECT_TYPE` | validate field semantics or derive object type from context |
| `WRONG_SEMANTIC_CLASS` | apply semantic filtering |
| `NO_SUITABLE_CONCEPT` | external Concept Scheme, manual review, vocabulary extension |

## 7. Candidate Discovery

`Candidate Discovery` ersetzt bewusst den früheren Begriff `Candidate Generation`.

Reconcilix erzeugt keine Zielbegriffe, sondern entdeckt mögliche `CandidateConcepts` in bestehenden kontrollierten Vokabularen oder angeschlossenen Concept Schemes.

Mögliche Verfahren:

- exact lookup
- lookup über alternative Terms
- broader / narrower / related concept traversal
- OpenRefine Reconciliation API
- xTree vocabulary lookup
- cross Concept Scheme lookup
- vector search, z. B. Qdrant
- embedding-based search, z. B. BGE
- LLM-assisted expansion oder disambiguation

Diese Verfahren sind keine Pipeline-Stufen. Sie sind mögliche Techniken innerhalb einer `Matching Strategy`.

## 8. Examples

### 8.1 ORTHOGRAPHIC_VARIANT: Altrarretabel

```text
SourceValue
  value = "Altrarretabel"

InterpretationGraph
  InterpretationNode[0]
    value = "Altrarretabel"
    level = 0

  ReconciliationResult[0]
    success = false
    pattern = ORTHOGRAPHIC_VARIANT

  InterpretationNode[1]
    parent = InterpretationNode[0]
    value = "Altarretabel"
    method = Orthographic Normalization

  ReconciliationResult[1]
    success = pending / evaluated
    pattern = COMPOSITE_TERM or SEMANTIC_GENERALIZATION
```

### 8.2 MULTI_CONCEPT_EXPRESSION: Altar-Kanzel-Orgelprospekt

```text
SourceValue
  value = "Altar-Kanzel-Orgelprospekt"

InterpretationGraph
  InterpretationNode[0]
    value = "Altar-Kanzel-Orgelprospekt"
    pattern = MULTI_CONCEPT_EXPRESSION

  InterpretationNode[1]
    value = "Altar"

  InterpretationNode[2]
    value = "Kanzel"

  InterpretationNode[3]
    value = "Orgelprospekt"
```

### 8.3 CONTEXT_DEPENDENT: Bogen in a sentence

```text
SourceValue
  value = "Der Herr von Bogen ging mit dem Bogen über den Bogen"

InterpretationGraph
  InterpretationNode[0]
    value = "Bogen"
    spanStart = ...
    spanEnd = ...
    pattern = CONTEXT_DEPENDENT

  InterpretationNode[1]
    value = "Bogen"
    spanStart = ...
    spanEnd = ...
    pattern = CONTEXT_DEPENDENT

  InterpretationNode[2]
    value = "Bogen"
    spanStart = ...
    spanEnd = ...
    pattern = CONTEXT_DEPENDENT
```

Jeder `InterpretationNode` kann eigene Kontextinformationen und eigene `CandidateConcepts` erhalten.

## 9. Relation to ME001

ME001 definiert die `Vocabulary Matching Patterns`. ME002 beschreibt, wie Reconcilix diese Patterns verwendet, um eine passende `Matching Strategy` auszuwählen.

ME001:

```text
What kind of matching situation was observed?
```

ME002:

```text
What should Reconcilix do next?
```

## 10. Open Questions

- Soll `VocabularyMatchingPattern` ausschließlich im `ReconciliationResult` gespeichert werden oder zusätzlich am `InterpretationNode` gespiegelt werden?
- Wie werden mehrere Patterns pro `InterpretationNode` technisch repräsentiert?
- Wie wird die höchste erfolgreiche Interpretation bestimmt?
- Wann wird eine weitere `Matching Strategy` abgebrochen?
- Wie werden Expert:innenentscheidungen zurück in den `InterpretationGraph` geschrieben?
