Skip to main content

Exporting the environment's content as a catalog

Semantic Treehouse (STH) is a vocabulary service in the terminology of dataspaces. The role of a vocabulary service is to facilitate semantic interoperability within and between data spaces and is described in the Data Spaces Support Center Blueprint. Semantic Treehouse is listed as a vocabulary service in the DSSC Toolbox. Because the core values of the dataspace movement include decentralisation, many vocabulary hubs may co-exist and their interoperability becomes essential.

STH supports this interoperability by offering a machine-readable export of the environment's content as a catalog. This is done primarily using mainly the Data Catalog Vocabulary standard (DCAT) supplemented by ontologies for vocabulary management, therefore we call it a DCAT+ export. The selection of DCAT follows the FAIR principles for making data Findable, Accessible, Interoperable, and Reusable. The expressiveness of DCAT makes the semantic specifications and all derived artifacts easily discoverable (also harvestable by tools) and reusable. The export allows users to find datasets as well. Of course, STH stores specifications, not datasets. But STH stores specifications for these datasets and stores the mappings that can help users discover datasets compatible with related specifications.

DCAT export as the base layer​

The DCAT standard was designed for datasets and not so much for the schemas to which datasets conform. Still it holds value, when we consider our specifications and their versions as Datasets with Distributions. The export describes the specifications and their versions and describes the different distributions for each, e.g. one 'Invoice v1.0' specification version may have separate distributions for its XML schema, JSON schema and SHACL constraints file. For each distribution, there is the explicit notion of dcat:landingPage, dcat:accessURL and dcat:downloadURL, as well as the format used to express the distribution (e.g., JSON Schema in JSON format, JSON Schema in YAML format). This information makes it much easier for tooling to find specific artifacts and to know how to automatically process them.

The additions: increasing findability and supporting mappings​

Another layer of information was added to support several use cases, and findability of resources and findability of validation and example artifacts. This includes the use case of mappings from one application profile / message model version to another. Consider the scenario where users are looking for datasets. These datasets conform to a specific specification version, and there is a limited number of datasets available. Mappings to compatible specification versions are provided by the catalog, and using these links the users can find other datasets that conform to mapped specification versions.

In order to achieve this, several additional ontologies are used in this export (ADMS, PROF, PMAP). We use ADMS for annotating versions (though version 2.0 of ADMS focused on re-using DCAT as much as possible some additions such as prev and next remain). We use PROF to describe the roles of distributions and specifications and to annotate artifacts with their role, such as 'example', 'constraints', 'schema', 'mapping'. Because of a lack of expressivity in the existing ontologies, we follow the suggestion in a recent research paper "The Vocabulary Hub as a Catalog for Semantic Artifacts for Discovery and Alignment of Datasets" presented at SDS 2026 conference, to add the PMAP ontology. The PMAP ontology makes it possible to express source and target profiles of mappings explicitly, which is crucial for finding related profiles obviously.

How to export the catalog​

The 'Specifications' overview page features a top-right download button. The offered file is expressed in Turtle format, with the .ttl extension. Note that only the specifications that the user has access to are included, and only the specification versions that are set to status 'Published'.