Skip to content

Ontologies

The Ontologies module is the semantic resource administration area of OpenSPAD.

It manages ontology discovery, analysis, import, Neo4j persistence, suite resolution, vectorization and semantic search support.

Current implementation status

The Ontologies module is substantially implemented.

It is not a placeholder page. The backend currently exposes real APIs for:

  • database status
  • ontology listing and detail retrieval
  • direct ontology import
  • ontology deletion
  • URL analysis before import
  • importing analyzed ontology sets
  • complete ontology suite analysis
  • complete suite import
  • vectorization setup
  • vectorization status
  • vectorization runs
  • vector search preview
  • use-case semantic search

Some search and analysis functionality is still exposed through test or preview endpoints and should therefore be treated as evolving functionality.


Architecture overview

flowchart LR
    URL[Ontology URL]
    ANALYZE[Analyze]
    SUITE[Resolve Suite]
    REVIEW[Review]
    IMPORT[Import]
    NEO[(Neo4j)]
    VECTOR[Vectorization]
    SEARCH[Semantic Search]

    URL --> ANALYZE
    ANALYZE --> SUITE
    SUITE --> REVIEW
    REVIEW --> IMPORT
    IMPORT --> NEO
    NEO --> VECTOR
    VECTOR --> SEARCH

The important OpenSPAD principle is:

Analyze first, import second.

Ontology discovery and suite resolution can therefore be inspected before data is written to Neo4j.


Administration page

The Ontologies administration page is available at:

/administration/ontologies/

The page is backed by real ontology administration logic and is connected to multiple API endpoints.


Current API map

Page

GET /ontologies/

Renders the Ontologies administration interface.


Database status

GET /api/ontologies/db-status

Used to inspect whether the ontology administration backend can access its database environment and ontology data store.


List ontologies

GET /api/ontologies

Returns the ontology records currently known to OpenSPAD.


Ontology details

GET /api/ontologies/{ontology_id}

Returns detailed information for one ontology.


Direct import

POST /api/ontologies/import

Imports ontology data into OpenSPAD.

This is a database-changing operation.


Delete ontology

DELETE /api/ontologies/{ontology_id}

Removes a selected ontology from the OpenSPAD ontology store.

Deletion is a database-changing operation.


Analyze before import

OpenSPAD provides a separate analysis stage before import.

POST /api/ontologies/analyze-url

This allows an ontology URL to be inspected before committing it to the database.

Conceptually:

flowchart LR
    URL[Ontology URL]
    FETCH[Fetch / Resolve]
    PARSE[Parse Metadata]
    GRAPH[Inspect Relations]
    RESULT[Analysis Result]

    URL --> FETCH --> PARSE --> GRAPH --> RESULT

The analysis stage is valuable because ontology sources can contain:

  • imports
  • reverse references
  • dependencies
  • extension modules
  • alignment ontologies
  • multiple RDF distributions

OpenSPAD can therefore inspect the semantic neighborhood before import.


Import analyzed ontologies

After analysis, OpenSPAD can import the reviewed result using:

POST /api/ontologies/import-analysis

This separates two concerns:

flowchart LR
    A[Analysis]
    R[Human / System Review]
    I[Import]
    DB[(Neo4j)]

    A --> R --> I --> DB

This is safer than automatically writing every discovered resource into the graph.


Complete suite analysis

OpenSPAD also exposes:

POST /api/ontologies/analyze-suite

The suite resolver is intended for ontology families rather than isolated ontology files.

The backend contains dedicated logic for resolving:

  • a core ontology
  • extension modules
  • dependencies
  • alignment documents
  • linked RDF distributions

Conceptually:

flowchart TB
    CORE[Core Ontology]
    EXT1[Extension A]
    EXT2[Extension B]
    DEP[Dependency]
    ALIGN[Alignment]

    CORE --> EXT1
    CORE --> EXT2
    CORE --> DEP
    CORE --> ALIGN

The exact content depends on the ontology suite being analyzed.


Import complete suite

After suite analysis, the resolved set can be imported using:

POST /api/ontologies/import-suite-analysis

Conceptually:

flowchart LR
    CORE[Core]
    EXT[Extensions]
    DEP[Dependencies]
    ALIGN[Alignments]
    IMPORT[Suite Import]
    NEO[(Neo4j)]

    CORE --> IMPORT
    EXT --> IMPORT
    DEP --> IMPORT
    ALIGN --> IMPORT
    IMPORT --> NEO

The distinction between suite analysis and suite import is intentional.


Ontology suite resolution

The backend contains dedicated resolver logic for RDF documents and ontology-suite resources.

Its responsibilities include determining which machine-readable ontology resource should be used when a public ontology page provides multiple links or distributions.

A simplified conceptual sequence is:

flowchart TD
    A[Ontology entry URL]
    B[Inspect resource]
    C[Find RDF candidates]
    D[Score candidates]
    E[Resolve preferred RDF document]
    F[Analyze imports and dependencies]

    A --> B --> C --> D --> E --> F

This allows OpenSPAD to work with ontology publication sites rather than requiring every user to manually find the final RDF/Turtle file.


DICON-related suite support

The current backend contains dedicated DICON resolution logic, including module and alignment resolution.

This indicates that the suite-analysis subsystem is not only generic: it also includes ontology-family-specific handling where required.

These resolver functions should be considered implementation-specific and may continue to evolve as ontology suites change their publication structure.


Neo4j persistence

Imported ontology information is persisted in the OpenSPAD graph database.

Conceptually:

flowchart TB
    O[Ontology]
    C1[Class]
    C2[Class]
    P1[Property]

    O --> C1
    O --> C2
    O --> P1

    C1 -->|semantic relation| C2

The graph structure allows OpenSPAD to retain relationships between semantic concepts rather than flattening ontology content into simple rows.


Ontology lifecycle

A simplified ontology lifecycle is:

stateDiagram-v2
    [*] --> Candidate
    Candidate --> Analyzed
    Analyzed --> Imported
    Imported --> Vectorized
    Vectorized --> Searchable
    Imported --> Deleted
    Vectorized --> Deleted

The exact stored status fields may differ internally, but this represents the functional workflow of the current module.


Vectorization

The ontology backend includes a complete vectorization subsystem.

Current API endpoints are:

POST /api/ontologies/vectorization/setup
GET  /api/ontologies/vectorization/status
POST /api/ontologies/vectorization/run
POST /api/ontologies/vectorization/search-preview

The vectorization subsystem creates semantic embeddings for ontology resources so that they can be searched by meaning rather than only exact labels.


Embedding model

The current OpenSPAD vectorization configuration uses FastEmbed with:

BAAI/bge-small-en-v1.5

and produces:

384-dimensional vectors

The intended similarity measure is cosine similarity.


Vector indexes

OpenSPAD maintains separate vector indexes for major ontology object types:

openspad_ontology_vectors
openspad_class_vectors
openspad_property_vectors

Conceptually:

flowchart TB
    MODEL[Embedding Model]

    ONT[Ontology Text]
    CLS[Class Text]
    PROP[Property Text]

    OIDX[(Ontology Vector Index)]
    CIDX[(Class Vector Index)]
    PIDX[(Property Vector Index)]

    MODEL --> ONT --> OIDX
    MODEL --> CLS --> CIDX
    MODEL --> PROP --> PIDX

Keeping separate indexes allows search behavior to distinguish ontology-level, class-level and property-level matches.


Text preparation

The vectorization backend contains dedicated text-building logic for:

  • ontology records
  • ontology classes
  • ontology properties

This is important because embeddings are only as useful as the semantic text supplied to the model.

Conceptually:

flowchart LR
    RAW[Ontology Metadata]
    BUILD[Semantic Text Builder]
    EMBED[Embedding Model]
    VECTOR[Vector]

    RAW --> BUILD --> EMBED --> VECTOR

Vectorization setup

POST /api/ontologies/vectorization/setup

The setup step prepares the vector-search infrastructure.

This includes creation or verification of the required vector indexes.


Vectorization status

GET /api/ontologies/vectorization/status

The status endpoint provides information about the current vectorization state and index state.

This allows the Administration UI or diagnostics to determine whether semantic search infrastructure is ready.


Vectorization run

POST /api/ontologies/vectorization/run

This performs vector generation and database updates.

The backend contains batch-processing logic, allowing large ontology datasets to be vectorized incrementally rather than as one monolithic operation.

Conceptually:

flowchart LR
    DATA[Ontology Data]
    BATCH[Batch]
    MODEL[Embedding Model]
    WRITE[Write vectors]
    INDEX[(Neo4j Vector Index)]

    DATA --> BATCH --> MODEL --> WRITE --> INDEX

Vector search preview

POST /api/ontologies/vectorization/search-preview

This endpoint allows semantic vector search to be tested before embedding the functionality into a final user workflow.

Because the endpoint is explicitly named search-preview, it should currently be treated as a preview / diagnostic capability.


Use-case semantic search

The backend also exposes:

POST /api/ontologies/usecase-search

This goes beyond simple nearest-neighbor vector search.

The implementation contains dedicated ranking logic for use-case searches and graph-based bonus calculations.

Conceptually:

flowchart TD
    Q[Use Case Text]
    EMB[Query Embedding]
    VEC[Vector Similarity]
    GRAPH[Graph Context]
    RANK[Weighted Ranking]
    RESULT[Ranked Ontology Concepts]

    Q --> EMB --> VEC
    VEC --> RANK
    GRAPH --> RANK
    RANK --> RESULT

This is an important architectural direction for OpenSPAD:

semantic similarity and graph structure can be combined rather than treated as separate systems.


Ontology Search Test page

The backend currently includes:

GET /ontology-search-test/

This is explicitly a test page.

It should be regarded as a development / evaluation surface rather than a final end-user interface.

The associated semantic search API is:

POST /api/ontologies/usecase-search

Functional status matrix

Capability Status
Ontologies administration page Implemented
Database status Implemented
List ontologies Implemented
Ontology details Implemented
Direct import Implemented
Delete ontology Implemented
Analyze URL before import Implemented
Import analyzed ontology set Implemented
Analyze complete suite Implemented
Import complete suite Implemented
RDF candidate resolution Implemented
Suite dependency resolution Implemented
Alignment resolution Implemented
DICON-specific resolution logic Implemented
Vector index setup Implemented
Vectorization status Implemented
Vectorization run Implemented
Vector search preview Implemented / preview
Use-case semantic search Implemented / evolving
Ontology Search Test page Development / test surface
Final end-user semantic search UX Evolving
Automatic ontology-to-pipeline decisions Target architecture

End-to-end ontology workflow

flowchart TD
    A[Enter Ontology URL]
    B[Analyze URL]
    C{Single resource<br/>or suite?}
    D[Analyze ontology]
    E[Analyze complete suite]
    F[Review discovered resources]
    G[Import]
    H[(Neo4j)]
    I[Vectorization Setup]
    J[Vectorization Run]
    K[Semantic Search]
    L[Use-Case Ranking]

    A --> B --> C
    C -->|single| D
    C -->|suite| E
    D --> F
    E --> F
    F --> G --> H
    H --> I --> J --> K --> L

Example

Assume an administrator provides a URL for a core ontology.

OpenSPAD can conceptually perform the following process:

1. Resolve the supplied URL
2. Identify the machine-readable RDF resource
3. Analyze ontology metadata
4. Discover imports / dependencies
5. Discover extensions and alignments
6. Present the analysis
7. Import the approved resources
8. Store semantic entities in Neo4j
9. Generate vector embeddings
10. Make ontology concepts available to semantic search

This pipeline is one of the central semantic foundations of OpenSPAD.


Why analysis and import are separate

Separating analysis from import has several advantages.

Transparency

The administrator can inspect what OpenSPAD discovered.

Safety

Unexpected external dependencies do not have to be imported automatically.

Reproducibility

The analyzed resource set can be treated as an explicit import decision.

Debugging

Problems in remote ontology publication structures can be diagnosed before graph data is changed.


Ontology data and semantic search

Ontology import and semantic search are linked but distinct stages.

flowchart LR
    RDF[RDF / OWL]
    GRAPH[Semantic Graph]
    TEXT[Search Text]
    VECTOR[Embeddings]
    QUERY[Use Case Query]
    MATCH[Ranked Concepts]

    RDF --> GRAPH
    GRAPH --> TEXT --> VECTOR
    QUERY --> MATCH
    VECTOR --> MATCH
    GRAPH --> MATCH

This hybrid architecture is important because pure vector similarity does not capture every semantic graph relationship, while graph traversal alone does not capture natural-language similarity.


Administration security

Ontology import, deletion and other database-changing operations should be treated as privileged administration functions.

OpenSPAD uses a server-side Administration mutation key for protected database-changing administration operations.

The key must not be committed to Git.


Current implementation vs future architecture

The Ontologies module is already one of the most mature OpenSPAD administration components.

The next architectural work is therefore less about creating a basic registry and more about integrating the existing semantic infrastructure with:

  • centralized schemas
  • plugin contracts
  • Pipeline Templates
  • the Pipeline Builder
  • project-specific Shared Semantic State
  • final IDS composition

Conceptually:

flowchart LR
    ONT[Ontology Knowledge]
    SEARCH[Semantic Search]
    ENRICH[Semantic Enrichment Plugin]
    STATE[Shared Semantic State]
    PIPE[Pipeline]
    IDS[IDS]

    ONT --> SEARCH --> ENRICH --> STATE --> PIPE --> IDS

Recommended next integration steps

flowchart TD
    A[Stable Ontology Import]
    B[Stable Vectorization]
    C[Semantic Search]
    D[Semantic Enrichment Plugin]
    E[Shared Semantic State]
    F[Pipeline Template Validation]
    G[Composer]
    H[IDS Output]

    A --> B --> C --> D --> E --> F --> G --> H

The first three layers already have substantial implementation.

The major remaining challenge is using those ontology results consistently inside the plugin pipeline.


Troubleshooting

Ontology URL cannot be analyzed

Possible causes include:

  • URL unavailable
  • HTML page does not expose a usable RDF distribution
  • RDF resource moved
  • external server blocks requests
  • unsupported or malformed RDF

Use the analysis stage before attempting import.

Suite analysis discovers unexpected resources

Review imports, dependencies and alignment resources before importing the suite.

Vectorization status is incomplete

Check whether vector indexes have been created and whether ontology/class/property records have been vectorized.

Search results are weak

Possible causes include:

  • ontology labels or descriptions provide little semantic context
  • relevant resources are not yet vectorized
  • the use case is too short or ambiguous
  • graph context does not yet connect the expected concepts

Ontology Search Test differs from final UX

That is expected. /ontology-search-test/ is explicitly a development/test surface.


Related documentation