Ontologies
The Ontologies module is the semantic resource administration area of OpenSPAD.
It manages ontology discovery, analysis, import, Neo4j persistence, suite resolution, vectorization and semantic search support.
Current implementation status
The Ontologies module is substantially implemented.
It is not a placeholder page. The backend currently exposes real APIs for:
- database status
- ontology listing and detail retrieval
- direct ontology import
- ontology deletion
- URL analysis before import
- importing analyzed ontology sets
- complete ontology suite analysis
- complete suite import
- vectorization setup
- vectorization status
- vectorization runs
- vector search preview
- use-case semantic search
Some search and analysis functionality is still exposed through test or preview endpoints and should therefore be treated as evolving functionality.
Architecture overview
flowchart LR
URL[Ontology URL]
ANALYZE[Analyze]
SUITE[Resolve Suite]
REVIEW[Review]
IMPORT[Import]
NEO[(Neo4j)]
VECTOR[Vectorization]
SEARCH[Semantic Search]
URL --> ANALYZE
ANALYZE --> SUITE
SUITE --> REVIEW
REVIEW --> IMPORT
IMPORT --> NEO
NEO --> VECTOR
VECTOR --> SEARCH
The important OpenSPAD principle is:
Analyze first, import second.
Ontology discovery and suite resolution can therefore be inspected before data is written to Neo4j.
Administration page
The Ontologies administration page is available at:
/administration/ontologies/
The page is backed by real ontology administration logic and is connected to multiple API endpoints.
Current API map
Page
GET /ontologies/
Renders the Ontologies administration interface.
Database status
GET /api/ontologies/db-status
Used to inspect whether the ontology administration backend can access its database environment and ontology data store.
List ontologies
GET /api/ontologies
Returns the ontology records currently known to OpenSPAD.
Ontology details
GET /api/ontologies/{ontology_id}
Returns detailed information for one ontology.
Direct import
POST /api/ontologies/import
Imports ontology data into OpenSPAD.
This is a database-changing operation.
Delete ontology
DELETE /api/ontologies/{ontology_id}
Removes a selected ontology from the OpenSPAD ontology store.
Deletion is a database-changing operation.
Analyze before import
OpenSPAD provides a separate analysis stage before import.
POST /api/ontologies/analyze-url
This allows an ontology URL to be inspected before committing it to the database.
Conceptually:
flowchart LR
URL[Ontology URL]
FETCH[Fetch / Resolve]
PARSE[Parse Metadata]
GRAPH[Inspect Relations]
RESULT[Analysis Result]
URL --> FETCH --> PARSE --> GRAPH --> RESULT
The analysis stage is valuable because ontology sources can contain:
- imports
- reverse references
- dependencies
- extension modules
- alignment ontologies
- multiple RDF distributions
OpenSPAD can therefore inspect the semantic neighborhood before import.
Import analyzed ontologies
After analysis, OpenSPAD can import the reviewed result using:
POST /api/ontologies/import-analysis
This separates two concerns:
flowchart LR
A[Analysis]
R[Human / System Review]
I[Import]
DB[(Neo4j)]
A --> R --> I --> DB
This is safer than automatically writing every discovered resource into the graph.
Complete suite analysis
OpenSPAD also exposes:
POST /api/ontologies/analyze-suite
The suite resolver is intended for ontology families rather than isolated ontology files.
The backend contains dedicated logic for resolving:
- a core ontology
- extension modules
- dependencies
- alignment documents
- linked RDF distributions
Conceptually:
flowchart TB
CORE[Core Ontology]
EXT1[Extension A]
EXT2[Extension B]
DEP[Dependency]
ALIGN[Alignment]
CORE --> EXT1
CORE --> EXT2
CORE --> DEP
CORE --> ALIGN
The exact content depends on the ontology suite being analyzed.
Import complete suite
After suite analysis, the resolved set can be imported using:
POST /api/ontologies/import-suite-analysis
Conceptually:
flowchart LR
CORE[Core]
EXT[Extensions]
DEP[Dependencies]
ALIGN[Alignments]
IMPORT[Suite Import]
NEO[(Neo4j)]
CORE --> IMPORT
EXT --> IMPORT
DEP --> IMPORT
ALIGN --> IMPORT
IMPORT --> NEO
The distinction between suite analysis and suite import is intentional.
Ontology suite resolution
The backend contains dedicated resolver logic for RDF documents and ontology-suite resources.
Its responsibilities include determining which machine-readable ontology resource should be used when a public ontology page provides multiple links or distributions.
A simplified conceptual sequence is:
flowchart TD
A[Ontology entry URL]
B[Inspect resource]
C[Find RDF candidates]
D[Score candidates]
E[Resolve preferred RDF document]
F[Analyze imports and dependencies]
A --> B --> C --> D --> E --> F
This allows OpenSPAD to work with ontology publication sites rather than requiring every user to manually find the final RDF/Turtle file.
DICON-related suite support
The current backend contains dedicated DICON resolution logic, including module and alignment resolution.
This indicates that the suite-analysis subsystem is not only generic: it also includes ontology-family-specific handling where required.
These resolver functions should be considered implementation-specific and may continue to evolve as ontology suites change their publication structure.
Neo4j persistence
Imported ontology information is persisted in the OpenSPAD graph database.
Conceptually:
flowchart TB
O[Ontology]
C1[Class]
C2[Class]
P1[Property]
O --> C1
O --> C2
O --> P1
C1 -->|semantic relation| C2
The graph structure allows OpenSPAD to retain relationships between semantic concepts rather than flattening ontology content into simple rows.
Ontology lifecycle
A simplified ontology lifecycle is:
stateDiagram-v2
[*] --> Candidate
Candidate --> Analyzed
Analyzed --> Imported
Imported --> Vectorized
Vectorized --> Searchable
Imported --> Deleted
Vectorized --> Deleted
The exact stored status fields may differ internally, but this represents the functional workflow of the current module.
Vectorization
The ontology backend includes a complete vectorization subsystem.
Current API endpoints are:
POST /api/ontologies/vectorization/setup
GET /api/ontologies/vectorization/status
POST /api/ontologies/vectorization/run
POST /api/ontologies/vectorization/search-preview
The vectorization subsystem creates semantic embeddings for ontology resources so that they can be searched by meaning rather than only exact labels.
Embedding model
The current OpenSPAD vectorization configuration uses FastEmbed with:
BAAI/bge-small-en-v1.5
and produces:
384-dimensional vectors
The intended similarity measure is cosine similarity.
Vector indexes
OpenSPAD maintains separate vector indexes for major ontology object types:
openspad_ontology_vectors
openspad_class_vectors
openspad_property_vectors
Conceptually:
flowchart TB
MODEL[Embedding Model]
ONT[Ontology Text]
CLS[Class Text]
PROP[Property Text]
OIDX[(Ontology Vector Index)]
CIDX[(Class Vector Index)]
PIDX[(Property Vector Index)]
MODEL --> ONT --> OIDX
MODEL --> CLS --> CIDX
MODEL --> PROP --> PIDX
Keeping separate indexes allows search behavior to distinguish ontology-level, class-level and property-level matches.
Text preparation
The vectorization backend contains dedicated text-building logic for:
- ontology records
- ontology classes
- ontology properties
This is important because embeddings are only as useful as the semantic text supplied to the model.
Conceptually:
flowchart LR
RAW[Ontology Metadata]
BUILD[Semantic Text Builder]
EMBED[Embedding Model]
VECTOR[Vector]
RAW --> BUILD --> EMBED --> VECTOR
Vectorization setup
POST /api/ontologies/vectorization/setup
The setup step prepares the vector-search infrastructure.
This includes creation or verification of the required vector indexes.
Vectorization status
GET /api/ontologies/vectorization/status
The status endpoint provides information about the current vectorization state and index state.
This allows the Administration UI or diagnostics to determine whether semantic search infrastructure is ready.
Vectorization run
POST /api/ontologies/vectorization/run
This performs vector generation and database updates.
The backend contains batch-processing logic, allowing large ontology datasets to be vectorized incrementally rather than as one monolithic operation.
Conceptually:
flowchart LR
DATA[Ontology Data]
BATCH[Batch]
MODEL[Embedding Model]
WRITE[Write vectors]
INDEX[(Neo4j Vector Index)]
DATA --> BATCH --> MODEL --> WRITE --> INDEX
Vector search preview
POST /api/ontologies/vectorization/search-preview
This endpoint allows semantic vector search to be tested before embedding the functionality into a final user workflow.
Because the endpoint is explicitly named search-preview, it should currently be treated as a preview / diagnostic capability.
Use-case semantic search
The backend also exposes:
POST /api/ontologies/usecase-search
This goes beyond simple nearest-neighbor vector search.
The implementation contains dedicated ranking logic for use-case searches and graph-based bonus calculations.
Conceptually:
flowchart TD
Q[Use Case Text]
EMB[Query Embedding]
VEC[Vector Similarity]
GRAPH[Graph Context]
RANK[Weighted Ranking]
RESULT[Ranked Ontology Concepts]
Q --> EMB --> VEC
VEC --> RANK
GRAPH --> RANK
RANK --> RESULT
This is an important architectural direction for OpenSPAD:
semantic similarity and graph structure can be combined rather than treated as separate systems.
Ontology Search Test page
The backend currently includes:
GET /ontology-search-test/
This is explicitly a test page.
It should be regarded as a development / evaluation surface rather than a final end-user interface.
The associated semantic search API is:
POST /api/ontologies/usecase-search
Functional status matrix
| Capability | Status |
|---|---|
| Ontologies administration page | Implemented |
| Database status | Implemented |
| List ontologies | Implemented |
| Ontology details | Implemented |
| Direct import | Implemented |
| Delete ontology | Implemented |
| Analyze URL before import | Implemented |
| Import analyzed ontology set | Implemented |
| Analyze complete suite | Implemented |
| Import complete suite | Implemented |
| RDF candidate resolution | Implemented |
| Suite dependency resolution | Implemented |
| Alignment resolution | Implemented |
| DICON-specific resolution logic | Implemented |
| Vector index setup | Implemented |
| Vectorization status | Implemented |
| Vectorization run | Implemented |
| Vector search preview | Implemented / preview |
| Use-case semantic search | Implemented / evolving |
| Ontology Search Test page | Development / test surface |
| Final end-user semantic search UX | Evolving |
| Automatic ontology-to-pipeline decisions | Target architecture |
End-to-end ontology workflow
flowchart TD
A[Enter Ontology URL]
B[Analyze URL]
C{Single resource<br/>or suite?}
D[Analyze ontology]
E[Analyze complete suite]
F[Review discovered resources]
G[Import]
H[(Neo4j)]
I[Vectorization Setup]
J[Vectorization Run]
K[Semantic Search]
L[Use-Case Ranking]
A --> B --> C
C -->|single| D
C -->|suite| E
D --> F
E --> F
F --> G --> H
H --> I --> J --> K --> L
Example
Assume an administrator provides a URL for a core ontology.
OpenSPAD can conceptually perform the following process:
1. Resolve the supplied URL
2. Identify the machine-readable RDF resource
3. Analyze ontology metadata
4. Discover imports / dependencies
5. Discover extensions and alignments
6. Present the analysis
7. Import the approved resources
8. Store semantic entities in Neo4j
9. Generate vector embeddings
10. Make ontology concepts available to semantic search
This pipeline is one of the central semantic foundations of OpenSPAD.
Why analysis and import are separate
Separating analysis from import has several advantages.
Transparency
The administrator can inspect what OpenSPAD discovered.
Safety
Unexpected external dependencies do not have to be imported automatically.
Reproducibility
The analyzed resource set can be treated as an explicit import decision.
Debugging
Problems in remote ontology publication structures can be diagnosed before graph data is changed.
Ontology data and semantic search
Ontology import and semantic search are linked but distinct stages.
flowchart LR
RDF[RDF / OWL]
GRAPH[Semantic Graph]
TEXT[Search Text]
VECTOR[Embeddings]
QUERY[Use Case Query]
MATCH[Ranked Concepts]
RDF --> GRAPH
GRAPH --> TEXT --> VECTOR
QUERY --> MATCH
VECTOR --> MATCH
GRAPH --> MATCH
This hybrid architecture is important because pure vector similarity does not capture every semantic graph relationship, while graph traversal alone does not capture natural-language similarity.
Administration security
Ontology import, deletion and other database-changing operations should be treated as privileged administration functions.
OpenSPAD uses a server-side Administration mutation key for protected database-changing administration operations.
The key must not be committed to Git.
Current implementation vs future architecture
The Ontologies module is already one of the most mature OpenSPAD administration components.
The next architectural work is therefore less about creating a basic registry and more about integrating the existing semantic infrastructure with:
- centralized schemas
- plugin contracts
- Pipeline Templates
- the Pipeline Builder
- project-specific Shared Semantic State
- final IDS composition
Conceptually:
flowchart LR
ONT[Ontology Knowledge]
SEARCH[Semantic Search]
ENRICH[Semantic Enrichment Plugin]
STATE[Shared Semantic State]
PIPE[Pipeline]
IDS[IDS]
ONT --> SEARCH --> ENRICH --> STATE --> PIPE --> IDS
Recommended next integration steps
flowchart TD
A[Stable Ontology Import]
B[Stable Vectorization]
C[Semantic Search]
D[Semantic Enrichment Plugin]
E[Shared Semantic State]
F[Pipeline Template Validation]
G[Composer]
H[IDS Output]
A --> B --> C --> D --> E --> F --> G --> H
The first three layers already have substantial implementation.
The major remaining challenge is using those ontology results consistently inside the plugin pipeline.
Troubleshooting
Ontology URL cannot be analyzed
Possible causes include:
- URL unavailable
- HTML page does not expose a usable RDF distribution
- RDF resource moved
- external server blocks requests
- unsupported or malformed RDF
Use the analysis stage before attempting import.
Suite analysis discovers unexpected resources
Review imports, dependencies and alignment resources before importing the suite.
Vectorization status is incomplete
Check whether vector indexes have been created and whether ontology/class/property records have been vectorized.
Search results are weak
Possible causes include:
- ontology labels or descriptions provide little semantic context
- relevant resources are not yet vectorized
- the use case is too short or ambiguous
- graph context does not yet connect the expected concepts
Ontology Search Test differs from final UX
That is expected. /ontology-search-test/ is explicitly a development/test surface.