# Vision

## xTree Reconciliation Framework

### Vision Statement

The Recocilix  Reconciliation Framework aims to become an open, extensible and vocabulary-aware framework for semantic entity resolution in cultural heritage data.

Rather than being tied to a single application or reconciliation protocol, the framework is designed as a reusable core that can be integrated into different workflows, user interfaces and research infrastructures.

OpenRefine is the first adapter, not the final application.

---

# Motivation

Cultural heritage data is heterogeneous.

Descriptions originate from different institutions, historical periods and documentation practices. They contain abbreviations, OCR artefacts, multilingual terminology, historical spellings and context-dependent expressions.

Examples include:

* `Kath. Dom St. Marien`
* `Matthiaskapelle`
* `Innenhof eines Kreuzgangs`
* `ehem. Pfalzkapelle`
* `Der Herr von Bogen geht mit dem Bogen über den Bogen.`

Traditional reconciliation systems treat these examples as strings.

Human experts, however, interpret them as references to entities embedded in linguistic and semantic context.

The framework therefore aims to bridge the gap between textual descriptions and structured knowledge systems.

---

# From String Matching to Entity Resolution

Traditional reconciliation follows a relatively simple model:

```text
search string
      ↓
lookup service
      ↓
identifier
```

This approach works well for controlled vocabularies and exact terminology.

However, many cultural heritage datasets require more than exact lookup.

The long-term goal of the framework is therefore to move from **string reconciliation** towards **entity resolution**.

Instead of processing isolated strings, the reconciliation core should eventually work on semantically enriched search requests.

Future requests may contain:

* the original text,
* detected mentions,
* document language,
* entity type hints,
* preferred vocabularies,
* contextual information,
* previously identified entities,
* user feedback.

Example:

```json
{
  "text": "Der Herr von Bogen geht mit dem Bogen über den Bogen.",
  "mentions": [
    {
      "text": "Herr von Bogen",
      "entityTypeHint": "Person"
    },
    {
      "text": "Bogen",
      "entityTypeHint": "Object"
    },
    {
      "text": "Bogen",
      "entityTypeHint": "Place"
    }
  ]
}
```

Instead of reconciling a single string, the framework would independently resolve each detected mention against one or more knowledge sources.

This transforms reconciliation into a semantic interpretation process.

---

# Human and Machine Collaboration

The framework is explicitly designed as a **human-in-the-loop** system.

Automatic candidate generation should support experts, not replace them.

Possible processing stages include:

* tokenisation,
* lemmatisation,
* stemming,
* morphological analysis,
* named entity recognition,
* abbreviation expansion,
* OCR correction,
* multilingual normalisation,
* terminology expansion,
* LLM-supported context interpretation.

Every stage contributes additional evidence for the final reconciliation.

The final decision, however, often remains a human task.

---

# Vocabulary-aware Reconciliation

Knowledge organisation systems are more than flat lists of preferred labels.

The framework therefore explicitly models:

* Vocabulary
* SubVocabulary
* Concept Scheme
* Source System

Different search strategies may be applied depending on the selected search domain.

Examples include:

* full vocabulary search,
* hierarchy-restricted search,
* terminology index,
* linguistic fallback,
* external authority lookup.

---

# OpenRefine is an Adapter

OpenRefine is currently the first user interface supported by the framework.

Its reconciliation protocol provides an excellent environment for interactive review and manual decision making.

Nevertheless, OpenRefine is only one possible adapter.

The same reconciliation core should later be usable from:

* Collection Management Systems
* REST APIs
* command line tools
* batch processing
* TEI/XML workflows
* OCR post-processing
* annotation environments
* future research infrastructures

This separation between adapter and reconciliation core is a fundamental architectural principle of the project.

---

# Architecture

The framework follows a package-oriented, hexagonal architecture.

```text
                Inbound Adapters

      OpenRefine
      REST API
      CLI
      xTree
      Batch Processing

               │
               ▼

        Reconciliation Core

   Tenant
   Vocabulary
   SubVocabulary
   SearchRequest
   Candidate
   SearchStrategy
   Ranking

               │
               ▼

          Source Systems / Clients

      xTree JSON
      xTree Solr
      xTree SPARQL
      Wikidata
      GND
      GBIF
      ...

               │
               ▼

        Output Adapters

      OpenRefine JSON
      REST JSON
      CSV
      internal APIs
```

The domain model should remain independent from the technical implementation of individual source systems.

---

# Design Principles

The project follows several core principles.

## Domain First

The domain model should describe concepts from knowledge organisation rather than implementation details.

Examples include:

* Vocabulary
* SubVocabulary
* Candidate
* Tenant
* SearchRequest
* SourceSystem

Technical concepts such as REST endpoints or Solr queries belong to infrastructure layers.

---

## Adapter-based Integration

External systems should communicate with the framework through adapters.

The reconciliation core should not know whether data originates from OpenRefine, digiCULT.web or another application.

---

## Tenant-aware Infrastructure

Different institutions require different vocabularies, credentials and technical infrastructures.

Therefore, every request is associated with a Tenant.

The Tenant controls:

* available vocabularies,
* available source systems,
* credentials,
* optional features.

---

## Technology-independent Source Systems

A Source System represents a knowledge source.

The technical access method is implemented by interchangeable clients.

Example:

```
Source System
    xTree

Clients
    xTree JSON
    xTree Solr
    xTree SPARQL
```

This allows migration to new technologies without changing the domain model.

---

## Open Source

The project is intended to remain framework-independent and suitable for publication under a permissive open-source licence such as MIT.

This facilitates:

* external contributions,
* academic collaboration,
* long-term sustainability,
* reuse in research infrastructures,
* integration into projects funded by organisations such as DFG or NFDI.

---

# Roadmap

## Short Term

* stable OpenRefine adapter
* hierarchy-aware SubVocabulary search
* SearchStrategy abstraction
* SubVocabularyTermIndex

## Medium Term

* linguistic preprocessing
* spaCy integration
* BGE (BAAI General Embedding) als Kandidatensucher
  *  Minds Mirakulix 2026-06-29: 
     *  Begriffe aus Zielvokabular als Embedding-Text erstellen 
     * -> mit bge-m3 in Vektor umwandeln 
     * -> nach Qdrant speichern
     * Quellstring mit Zusatzinfos über BGE in Vektor umwandeln
       * -> Suche in Qdrant
* configurable fallback strategies
* Solr client
* SPARQL/QLever client

## Long Term

Build a reusable reconciliation framework that supports semantic entity resolution across heterogeneous cultural heritage vocabularies while remaining independent from any single user interface or technical infrastructure.

The ultimate objective is not to reconcile strings.

The objective is to reconcile knowledge.
