# Architecture

## Purpose

The xTree Reconciliation Service v2 is evolving from a single OpenRefine endpoint into a framework for authority matching, vocabulary reconciliation and candidate generation.

OpenRefine is currently the first adapter, not the core of the architecture.

The central idea is:

```text
any client
  ↓
adapter
  ↓
reconciliation core
  ↓
source systems like xTree, Wikidata, ....
  ↓
candidates
```

## Architectural Style

The project follows a package-oriented, hexagonal architecture.

That means:

- the domain core should not depend on OpenRefine,
- OpenRefine is one inbound adapter,
- xTree JSON, xTree Solr, xTree SPARQL, Wikidata, GND and GBIF are outbound adapters,
- configuration and credentials are separated from code,
- tenant-specific access is resolved before clients are instantiated.

## Layers

### 1. Inbound adapters

Inbound adapters translate external requests into internal reconciliation requests.

Currently implemented:

- OpenRefine Reconciliation API

Potential future adapters:

- integration into collection managenment software systems
- REST API
- CLI tool
- batch processing
- TEI processing pipeline
- internal project-specific UIs

### 2. Reconciliation core

The core should contain domain concepts and orchestration logic.

Key domain objects:

- Tenant
- Vocabulary
- SubVocabulary
- Candidate
- SearchRequest
- SearchStrategy

The core should not know whether a request came from OpenRefine, a CLI command or another application.

### 3. Source systems

A source system describes where data lives.

Examples:

- xtree
- wikidata
- gnd
- gbif

A source system can have multiple technical clients.

Example xTree:

- xtree-json
- xtree-solr
- xtree-sparql

### 4. Technical clients

Clients implement the concrete access technology.

Examples:

- XtreeJsonApiClient
- XtreeSolrClient
- XtreeSparqlClient
- WikidataReconcileClient
- GndLobidClient
- GbifRestClient

### 5. Output adapters

Output adapters translate internal candidates into a target format.

Currently:

- OpenRefine JSON response

Future:

- JSON API response
- internal digiCULT.web response
- CSV/TSV export
- debugging format

## Current Request Flow

```text
OpenRefine POST
  ↓
reconcile.php
  ↓
ApiKeyAuthenticator
  ↓
TenantContext
  ↓
VocabularyRegistry
  ↓
SourceSystemFactory
  ↓
CredentialStore
  ↓
XtreeJsonApiClient
  ↓
xTree JSON API
  ↓
XtreeJsonApiResultMapper
  ↓
Candidate[]
  ↓
OpenRefineResponseBuilder
  ↓
OpenRefine JSON
```

## Vocabulary and SubVocabulary

### Vocabulary

A Vocabulary represents a concept scheme or authority source.

Example:

```php
new Vocabulary(
    id: 'http://matcult-the.vocnet.org',
    label: 'MCT',
    conceptSchemeId: 'http://matcult-the.vocnet.org',
    sourceSystem: 'xtree'
);
```

### SubVocabulary

A SubVocabulary represents a defined entry point or search domain inside a Vocabulary.

Example:

```php
new SubVocabulary(
    id: 'http://matcult-the.vocnet.org/00000019',
    label: 'MCT: Objektfacette',
    conceptSchemeId: 'http://matcult-the.vocnet.org'
);
```

OpenRefine sees both Vocabulary and SubVocabulary as `defaultTypes`.

Internally, the service resolves whether an OpenRefine type is a Vocabulary or SubVocabulary.

## Tenants and Credentials

A tenant is resolved from the API key.

The tenant defines:

- allowed vocabularies,
- allowed sub-vocabularies,
- allowed source systems,
- feature flags.

Credentials are resolved by tenant and technical client.

Example:

```php
return [
    'tenant-id-XYZ' => [
        'xtree-json' => [
            'username' => '...',
            'password' => '...',
        ],
    ],
];
```

Credentials must not be committed to Git.

## SourceSystemFactory

The SourceSystemFactory creates the correct client for a Vocabulary and Tenant.

It uses:

- Vocabulary.sourceSystem
- providers.php
- default_client
- CredentialStore

It should not contain:

- ranking logic,
- mapping logic,
- fallback logic,
- OpenRefine-specific logic.

## Search Strategies

The next architectural step should be the introduction of SearchStrategy.

Expected strategies:

```text
VocabularySearchStrategy
  → getSearchVocItemsByTerm()

SubVocabularySearchStrategy
  → getFetchHierarchy()

TermIndexFallbackStrategy
  → SubVocabularyTermIndex

LinguisticFallbackStrategy
  → spaCy / NLP
```

This will keep the reconciliation core independent from the details of xTree search methods.

## OpenRefine as Adapter

OpenRefine is not the main application.

It is the first integration target.

This means the POST method should eventually be split into:

```text
OpenRefine request parser
  ↓
internal SearchRequest
  ↓
Reconciliation core
  ↓
Candidate[]
  ↓
OpenRefine response builder
```

The same core could then be used by digiCULT.web or other systems.

## Current Technical Debt

### reconcile.php is too large

It currently contains too many responsibilities:

- request logging,
- authentication,
- registry lookup,
- source system resolution,
- client creation,
- search call,
- mapping,
- OpenRefine response.

This should be moved into smaller controller and service classes.

### SubVocabulary is not yet active

SubVocabulary is already recognized and logged, but the search still uses `getSearchVocItemsByTerm()` on the full ConceptScheme.

Next step:

- implement `getFetchHierarchy()`,
- use SubVocabulary.id as node id,
- map hierarchy response into Candidate[].

### Fallbacks are provisional

Fallbacks currently live in providers.php and are marked for reconsideration in v3.

Possible future structure:

- config/fallbacks.php
- FallbackProfile
- fallback profile referenced by Vocabulary or SubVocabulary

## Architectural Direction

The long-term direction is a framework-like, package-oriented architecture:

```text
reconciliation-core
openrefine-adapter
xtree-json-client
xtree-solr-client
xtree-sparql-client
wikidata-client
gbif-client
gnd-client
```

The project should remain framework-independent and be usable from:

- plain PHP,
- Symfony,
- Laravel,
- CLI tools,
- xTree,
- Java-Applications,
- batch workflows.
