Ja. Ich würde unseren aktuellen Arbeitsstand so zusammenfassen:

## Discovery Profiles – konsolidierter Stand v0.1

Für den aktuellen Dehio-Scope arbeiten wir pragmatisch mit **IG = 1 pro SourceValue**. Das ist eine KISS-Laufzeitkonvention, keine Einschränkung des Domänenmodells. Compound Terms und Fälle mit mehreren InterpretationGraphs werden architektonisch mitgedacht, aber noch nicht umgesetzt.

Der Ablauf soll künftig grundsätzlich so gedacht werden:

```text
SourceValue
  ↓
InterpretationGraph
  ↓
InterpretationNode
  ↓
Discovery Profile
  ↓
Discovery Step(s)
  ↓
Adapter / Provider
  ↓
CandidateItems
```

Ein **Discovery Profile** ist eine geordnete Folge von Discovery Steps. Die Steps dürfen unterschiedliche Methodenfamilien kombinieren. Genau diese Durchmischung ist zentral.

Ein Discovery Step hat als Kernstruktur:

```text
DiscoveryStep
├── family
├── method
├── resource      [optional]
├── options       [optional]
└── condition
```

Zusätzlich gehört der Input dazu:

```text
INPUT
├── SOURCE_VALUE
│   ├── VALUE
│   ├── [LANG]
│   └── PREF_ENTITY_TYPE
│
└── CONTEXT
    ├── ROLE
    ├── TYPE
    └── VALUE
```

Die drei aktuellen Dehio-Situationen sind:

```text
Qdrant
= SourceValue + TEI-Kontext

RxStore / xTree
= SourceValue + Qualifier-Kontext

Lobid
= SourceValue
```

Dabei bleibt wichtig: `Text Context` ist der semantische Context Type; **TEI ist die Repräsentation des Context Values**, nicht der Context Type selbst.

### Methodenfamilien

Der aktuelle Arbeitsstand ist:

```text
DiscoveryMethod
│
├── LexicalDiscoveryMethod
│   ├── EXACT_STRING
│   ├── RIGHT_TRUNCATED
│   ├── LEFT_TRUNCATED
│   ├── BOTH_TRUNCATED
│   ├── LEVENSHTEIN
│   └── ...
│
├── MappingDiscoveryMethod
│   └── MAPPED_VOCABULARY
│
├── HierarchicalDiscoveryMethod
│   └── ...
│
└── SemanticDiscoveryMethod
    └── VECTOR_SIMILARITY
```

`*_TRUNCATED` bevorzugen wir gegenüber Prefix-/Suffix-Bezeichnungen, weil diese sprachlich missverständlich sein können.

`LEVENSHTEIN` ist ein Subtyp von `LexicalDiscoveryMethod`. Methodenspezifische Parameter gehören in `options`, z. B.:

```text
method: LEVENSHTEIN
options:
  max_distance: 2
```

### Conditions

Für v0.1 reichen zunächst:

```text
ALWAYS
ON_ZERO_RESULTS
```

Das Modell bleibt offen für spätere Bedingungen wie etwa:

```text
ON_RESULT_COUNT_BELOW
ON_SCORE_BELOW
ON_NO_ACCEPTABLE_RESULT
```

`ON_ZERO_RESULTS` soll sich auf das bisherige relevante Ergebnis des Profils beziehen, nicht nur mechanisch auf den unmittelbar vorherigen Step.

### Profil A – rein lexikalische Eskalation

Beispiel:

```text
Step 1
  family: LexicalDiscoveryMethod
  method: EXACT_STRING
  condition: ALWAYS

Step 2
  family: LexicalDiscoveryMethod
  method: RIGHT_TRUNCATED
  condition: ON_ZERO_RESULTS

Step 3
  family: LexicalDiscoveryMethod
  method: LEFT_TRUNCATED
  condition: ON_ZERO_RESULTS

Step 4
  family: LexicalDiscoveryMethod
  method: BOTH_TRUNCATED
  condition: ON_ZERO_RESULTS
```

### Profil B – hybride Strategy

Für Dehio aktuell besonders wichtig:

```text
Step 1
  family: LexicalDiscoveryMethod
  method: EXACT_STRING
  condition: ALWAYS

Step 2
  family: MappingDiscoveryMethod
  method: MAPPED_VOCABULARY
  resource: Dehio-Hilfsvokabular
  condition: ON_ZERO_RESULTS
```

Das ist bewusst eine **Durchmischung verschiedener Methodenfamilien**. Genau das soll ein Discovery Profile leisten.

Andere Profile können z. B. kombinieren:

```text
LexicalDiscoveryMethod
  +
HierarchicalDiscoveryMethod
```

### Method und Resource sind getrennt

Bei `MAPPED_VOCABULARY` ist die Trennung besonders wichtig.

Die Method ist:

```text
MAPPED_VOCABULARY
```

Die fachliche Ressource ist z. B.:

```text
Dehio-Hilfsvokabular
```

Das erlaubt später andere Ressourcen mit derselben Method.

Fachlich gilt dabei ausdrücklich:

```text
Mapping
= intellektueller Prozess

Matching / Discovery
= automatischer Prozess
```

Rx erzeugt also während der Discovery **kein Mapping**, sondern nutzt ein bereits intellektuell erzeugtes Mapping als Ressource.

Auch wenn künftig LLMs oder Lernschleifen Mappingvorschläge erzeugen, bleibt das Mapping selbst eine fachliche Wissensressource. Eine spätere maschinelle Unterstützung ändert diese Trennung nicht.

### Adapter und API

Die Discovery Method soll providerunabhängig sein.

Beispiel:

```text
method = EXACT_STRING
```

kann durch mehrere Adapter ausgeführt werden:

```text
RxStore adapter
  exactString()

xTree adapter
  exactString()
```

Wenn wir die xTree-API selbst kontrollieren, soll dort ebenfalls die neue Terminologie verwendet werden:

```text
searchType = exactString
```

Historische Begriffe wie:

```text
simpleString
simpleString*
```

würden dann nur noch als Legacy-Bezeichnungen bzw. Rückwärtskompatibilität behandelt.

Damit ergibt sich:

```text
1. Discovery Method Family
   LexicalDiscoveryMethod

2. Discovery Method
   EXACT_STRING

3. Adapter Operation
   exactString()

4. Provider/API
   searchType = exactString
```

### Architekturprinzip

Der zentrale Gedanke lautet damit:

> **Discovery Profiles beschreiben fachlich und providerunabhängig, welche Methoden in welcher Reihenfolge, mit welchen Inputs, Ressourcen, Optionen und Bedingungen ausgeführt werden.**

Adapter setzen diese Methoden technisch gegen RxStore, xTree, Qdrant, Lobid oder zukünftige Provider um.

Und für unser späteres kontrolliertes Vokabular zeichnet sich derzeit dieser Kern ab:

```text
Family
→ Method
→ [Resource]
→ [Options]
→ Condition
```

