User Tools

Site Tools


wiki:1_road-to-fair-strategy:start

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
wiki:1_road-to-fair-strategy:start [2026/08/06 00:05] esoedingwiki:1_road-to-fair-strategy:start [2026/08/19 12:42] (current) dkottmeier
Line 1: Line 1:
 +====== The Road-to-FAIR Strategy ======
 +
 +The Road-to-FAIR Strategy provides the implementation framework for realizing the HMC vision described above. As discussed in the [[:preface|Preface]], the FAIR Guiding Principles define desired properties but do not prescribe the standards, technologies, responsibilities, or organizational processes through which these properties should be achieved.
 +
 +The Road-to-FAIR Strategy translates these general objectives into coordinated and actionable measures. It provides a common structure through which technical, semantic, organizational, and procedural activities can be related to one another, prioritized, and evaluated. This is particularly important because research data management encompasses a broad range of interdependent topics, including metadata quality, semantic interoperability, institutional responsibilities, technical infrastructures, and readiness for emerging forms of data-intensive research.
 +
 +The strategy addresses four closely connected areas. It:
 +
 +  * defines and prioritizes implementation objectives through a common set of **FAIR building blocks
 +  * identifies the stakeholder groups involved in research data management and clarifies their respective responsibilities
 +  * establishes workflows through which relevant information can be created, maintained, and transferred between responsible actors and systems, and
 +  * specifies the technical data ecosystem required to support these responsibilities and workflows.
 +
 +The Road-to-FAIR Strategy consequently treats research data management as a distributed and coordinated activity. Scientific information is generated at different stages of the research process and by a wide range of actors, many of whom may not primarily identify themselves as participants in research data management. The strategy seeks to ensure that this information is captured close to its point of origin, maintained by appropriate authoritative sources, and made available for repositories, aggregation services, scientific models, simulations, and machine-assisted analysis.
 +
 +==== 1. Defining the FAIR Building Blocks ====
 +
 +The [[..:2_fair-building-blocks:|FAIR building blocks]] define the principal implementation objectives of the Road-to-FAIR Strategy. They provide a structured means of assessing which elements of FAIR are already supported and where further coordinated action is required.
 +
 +Within the Helmholtz Research Field Earth and Environment, Findability and Accessibility are comparatively well supported for data deposited in established institutional or disciplinary repositories. Repository registries and discovery services, including re3data, FAIRsharing, and OpenAIRE, assist users in identifying appropriate repositories and locating relevant datasets. Open-science policies further support access to research data, while the provision of metadata for restricted datasets enables their discovery and communicates the conditions under which access may be obtained.
 +
 +The principal remaining challenges concern the **Interoperability** and **Reusability** of heterogeneous research data. Repositories, projects, research programmes, disciplines, and observational networks frequently apply different standards and procedures when describing their data. As a result, information may be difficult to combine across repositories and, in some cases, even within individual infrastructures. Projects that integrate heterogeneous datasets must therefore devote substantial effort to data cleaning, semantic harmonization, metadata enrichment, and the reconciliation of incompatible structures.
 +
 +The Road-to-FAIR [[..:2_fair-building-blocks:|building blocks]] define practical measures intended to reduce these barriers. Measures supporting Interoperability include:
 +
 +  * Common and agreed procedures for referencing systematically described entities through the use of persistent identifiers (PIDs)
 +  * A shared understanding across systems and disciplines through the coordinated use of semantic artefacts
 +  * Common methods for structuring and exchanging data and metadata through the application of interfaces, protocols, exchange formats, and schemas
 +  * Consistent approaches to data dissemination through self-describing data packages, such as FAIR Digital Objects and RO-Crates.
 +
 +Measures supporting Reusability include the systematic ddocumentation of provenance, data quality, access and use constraints, and other contextual information required to assess whether data are suitable for a particular purpose.
 +
 +The building blocks do not constitute a single technical specification. Rather, they define areas in which institutions, communities, and infrastructures must establish and implement shared agreements.
 +
 +==== 2. Defining Stakeholders and Responsibilities ==== 
 +
 +The Road-to-FAIR Strategy regards research data management as a distributed institutional responsibility. It cannot be assigned exclusively to individual researchers, data managers, or repositories.
 +
 +The information required to describe a research dataset is created and maintained by different stakeholder groups. These may include researchers, technicians, laboratory and field personnel, administrative units, libraries, data stewards, repository operators, infrastructure providers, and organizational management. No single stakeholder normally possesses all information required to produce a complete and accurate description of a digital research object.
 +
 +Responsibilities should therefore be assigned to those actors or systems best positioned to create, verify, and maintain the relevant information. Institutions must communicate these expectations clearly and, where appropriate, formalize them through policies, role descriptions, and agreed procedures. The identification of stakeholder groups and their responsibilities thus provides the organizational foundation for subsequent implementation.
 +
 +==== 3. Establishing Coordinated Workflows ==== 
 +
 +Distributed responsibilities require coordinated workflows. The Road-to-FAIR Strategy therefore defines procedures through which information is captured close to its point of origin, maintained by an appropriate authoritative source, and transferred to downstream information systems.
 +
 +This approach differs from workflows in which researchers are asked to reconstruct all required metadata only at the point of data publication. Information concerning people, organizations, projects, instruments, samples, methods, licences, or administrative conditions may already exist in dedicated systems and should, wherever possible, be obtained from these authoritative sources.
 +
 +Relevant information should therefore move continuously through interfaces connecting research processes, administrative systems, technical services, and data infrastructures. Such workflows reduce redundant data entry, improve consistency, and distribute the effort of metadata creation across the research data lifecycle.
 +
 +Their implementation requires both institutional support and technical services capable of exchanging, validating, aggregating, and enriching metadata. Workflows must also define how information is updated, who is responsible for correcting errors, and how changes are propagated across connected systems.
 +
 +==== 4. Providing an Interoperable Data Ecosystem ==== 
 +
 +The final component of the Road-to-FAIR Strategy is the definition and provision of a coordinated technical data ecosystem. This ecosystem comprises the tools, services, registries, interfaces, and infrastructures required to support the responsibilities and workflows established in the preceding steps.
 +
 +Relevant components may include persistent identifier services, electronic laboratory and field notebooks, sample and instrument management systems, institutional information systems, metadata editors, vocabulary services, validation tools, repositories, interfaces, and aggregation services.
 +
 +These components should not operate as isolated applications. They must form an interoperable environment in which data and metadata can be exchanged, validated, enriched, and reused across organizational and disciplinary boundaries. The technical ecosystem thereby enables information to be captured at its point of origin, maintained by responsible actors or authoritative systems, and transferred reliably to repositories and other downstream services.
 +
 +The data ecosystem provides the operational foundation through which assigned responsibilities and agreed workflows are translated into sustainable practice.
 +
 +==== Operationalizing FAIR ==== 
 +
 +The Road-to-FAIR Strategy provides the implementation framework through which the FAIR objectives described in the Preface are translated into coordinated practice. Its distinctive contribution is the integration of FAIR building blocks, stakeholder responsibilities, coordinated workflows, and an interoperable data ecosystem into a common implementation approach.
 +
 +Implementation should be understood as incremental rather than binary. Rich metadata alone is insufficient: FAIRness also depends on semantic alignment, interoperable structures, and clearly documented conditions for access and reuse. As emphasized in the Preface, FAIR does not imply unrestricted open access; data may remain subject to legal, ethical, contractual, or institutional access controls when their metadata, access conditions, and procedures are adequately documented.
 +
 +Taken together, these elements establish the socio-technical conditions required to operationalize FAIR. By connecting strategic objectives with community agreements, institutional responsibilities, coordinated workflows, and an interoperable data ecosystem, the strategy provides a structured pathway from a general commitment to FAIR towards a progressively harmonized and reusable Helmholtz research data space.
 +
 +/*
 +
 ====== Road-to-FAIR-Strategy ====== ====== Road-to-FAIR-Strategy ======
  
Line 11: Line 83:
 The RTF strategy thus sets the beginning and end of all our activities. It starts with defining RDM as a community task, orchestrated by the RDM concepts and personnel, where many people, often beyond their current awareness, play a role in documenting science, creating knowledge, and ultimately allowing for next-level science in the form of AI and ML applications. It finally suggests ways how to make this knowledge accessible and usable. The strategy supports the various types of information in where they are created, along and across the documentation process, towards their intermediate location in the repositories, ready for aggregation and enrichment processes, on their way to serve as fuel for system understanding, scientific models and simulations, and ultimately decision making for a better future of mankind served by the harmonized FAIR metadata space. The RTF strategy thus sets the beginning and end of all our activities. It starts with defining RDM as a community task, orchestrated by the RDM concepts and personnel, where many people, often beyond their current awareness, play a role in documenting science, creating knowledge, and ultimately allowing for next-level science in the form of AI and ML applications. It finally suggests ways how to make this knowledge accessible and usable. The strategy supports the various types of information in where they are created, along and across the documentation process, towards their intermediate location in the repositories, ready for aggregation and enrichment processes, on their way to serve as fuel for system understanding, scientific models and simulations, and ultimately decision making for a better future of mankind served by the harmonized FAIR metadata space.
  
-1. The beginning - the FAIR building blocks+==== 1. The beginning - the FAIR building blocks ====
  
 The FAIR building blocks define the core and the some immediate goals on our road to FAIR. The building blocks are aligned along the FAIR princpiles, as we ask ourselves: where do we stand in the implementation of FAIR? In Helmholtz we assume, that a significant part of the FAIR principles are already solved and implemented. These parts comprise Findability and Accesability for all data, that has been deposited in well managed institutional or disciplinary repositories. This information is relatively easy to find, e.g. through meta-databases like re3data, fairsharing or OpenAire tools. The metadata and often the data itself can typically be accessed, as Helmholtz follows an open science policy, and strives to publish its data whereever possible. The big challenges within our organization are the Interoperability and the Reusability of our data. All repositories, projects, research programs, disciplines, networks follow own rules, standards, and procedures in describing their data. This makes it nearly impossible to aggregate information from more than one repository, or even often within repositories. As a consequence data reuse on very heterogeneous datasets is very hard to conduct. In fact projects, who try to model on heterogeneous data sets and aggreagate information, spend by far most of their time, to clean data, harmonize semantic expressions, and enrich metadata with missing information, in order to compile useful, high-quality data sets.  The FAIR building blocks define the core and the some immediate goals on our road to FAIR. The building blocks are aligned along the FAIR princpiles, as we ask ourselves: where do we stand in the implementation of FAIR? In Helmholtz we assume, that a significant part of the FAIR principles are already solved and implemented. These parts comprise Findability and Accesability for all data, that has been deposited in well managed institutional or disciplinary repositories. This information is relatively easy to find, e.g. through meta-databases like re3data, fairsharing or OpenAire tools. The metadata and often the data itself can typically be accessed, as Helmholtz follows an open science policy, and strives to publish its data whereever possible. The big challenges within our organization are the Interoperability and the Reusability of our data. All repositories, projects, research programs, disciplines, networks follow own rules, standards, and procedures in describing their data. This makes it nearly impossible to aggregate information from more than one repository, or even often within repositories. As a consequence data reuse on very heterogeneous datasets is very hard to conduct. In fact projects, who try to model on heterogeneous data sets and aggreagate information, spend by far most of their time, to clean data, harmonize semantic expressions, and enrich metadata with missing information, in order to compile useful, high-quality data sets. 
  
-2. The Core concepts: Who is responsible for what or defining stakeholder groups and their roles in our organisation+The FAIR building blocks specify the concrete measures needed to make research data usable across infrastructure. The suggested measures to achieve interopeability are 1. the consequent use of persistent identifiers for redundant and recurring information, 2. applying agreed upon semantic concepts to metadata where applicable, 3. agreeing on standardized interfaces, protocols and formats to exchange metadata, 4. defining standards to expose our metadata and data in a uniform machine readable way. The suggested measures to achieve reusability are recording provenance and license information. They serve as practical targets for institutions and repositories, showing which technical and organizational components must be implemented to improve the findability, accessibility, interoperability, and reusability of data. 
 + 
 +==== 2. The Core concepts: Who is responsible for what or defining stakeholder groups and their roles in our organisation ==== 
 + 
 + 
 +The Road-to-FAIR Strategy treats research data management as a distributed institutional responsibility rather than an activity that can be assigned exclusively to researchers or repositories. The information required to describe a dataset is created and maintained by different stakeholder groups, including researchers, technicians, administrative units, organizational management, libraries, data stewards, and repository operators. No single stakeholder normally possesses all information required for a complete and accurate description. The organization should therefore actively express expectations towards those stakeholder groups and document them through internal policies.  
 + 
 +==== 3. Activating the community: Defining workflows and standard procedures to capture core information and pass it on within the system ==== 
 + 
 + 
 +To coordinate these distributed responsibilities, the strategy defines workflows through which information is captured close to its point of origin, maintained by an appropriate authoritative source, and transferred to downstream data systems. This approach replaces the common practice of requesting all metadata from researchers only at the time of data publication. Instead, relevant information should flow continuously through institutional information interfaces connecting administrative systems, research processes, technical services, and data infrastructures. Its implementation requires both institutional support and technical services capable of exchanging, validating, aggregating, and enriching metadata. 
 + 
 +==== 4. Making it possible: Enabling the community by defining and providing the technical **data ecosystem** supporting the data ==== 
 + 
 +In a fourth step, the Road-to-FAIR Strategy requires the definition, construction, and activation of a coordinated data ecosystem. This ecosystem comprises the services, tools, registries, interfaces, and technical infrastructures needed to support the stakeholder responsibilities and workflows defined in the preceding steps. Its purpose is to enable information to be captured at its point of origin, maintained by the responsible actors or authoritative systems, and transferred reliably to repositories and other downstream services. Relevant components may include persistent identifier registries, electronic laboratory and field notebooks, sample and instrument management systems, institutional information systems, vocabulary services, metadata editors, validation tools, interfaces, and aggregation services. These components must not operate as isolated applications, but as an interoperable environment in which metadata can be exchanged, validated, enriched, and reused across organizational and disciplinary boundaries. The data ecosystem therefore provides the technical foundation that translates assigned responsibilities and agreed workflows into sustainable operational practice.
  
-A complete and comprehensive description of a dataset is not trivial. It requires in-depth knowledge about techniques how to describe the different actorsorganizationsmeasured parametersinstruments and methods used to create the dataset and many other aspects needed to describe the circumstances under which a dataset or sample was produced or acquired. Experience showsthat not one person can provide all this information herself. We therefore suggest to adopt a community approach in collecting information related to the description of datasets, and take a look, who in our organisation assumes which role and what data the people assuming this role are possible administrating. We thus have to take a look at the organisation itself and it's stakeholdersAt Helmholtz we have identified +The Road-to-FAIR Strategy complements the FAIR principles by translating their high-level objectives into a coordinated implementation framework. While FAIR defines the desired properties of research data and metadatait does not prescribe which standards, technologiesresponsibilitiesor processes should be used. The Road-to-FAIR Strategy addresses this gap by defining concrete building blocks, assigning responsibilities to relevant stakeholder groupsestablishing workflows for the creation and maintenance of metadata, and specifying the technical services and interfaces required to support these processes.
  
-3. Activating the community: Defining workflows and standard procedures to capture cor information and pass it on within the system+A central element of the strategy is the development of shared community agreements. Persistent identifiers, metadata schemas, semantic vocabularies, exchange protocols, provenance models, licenses, and quality procedures must be selected and applied consistently across infrastructures. The strategy therefore treats FAIR implementation as a collaborative and incremental process rather than a binary state. Institutions can identify areas that are already comparatively mature, prioritize remaining gaps, and progressively improve interoperability, reusability, and machine actionability.
  
-4Making it possible: Enabling the community by defining and providing the technical **data ecosystem** supporting the data+The strategy also recognizes that rich metadata alone are insufficientMetadata must use harmonized semantics, explicit relationships, and interoperable formats so that information can be reliably interpreted, combined, and reused across organizational and disciplinary boundaries. At the same time, FAIRness is distinguished from unrestricted openness: data may remain access-controlled where legal, ethical, or contractual requirements apply, provided that access conditions and procedures are clearly documented.
  
 +Overall, the Road-to-FAIR Strategy establishes the socio-technical conditions needed to operationalize FAIR. It connects strategic goals with community agreements, institutional responsibilities, practical workflows, and an interoperable data ecosystem. In doing so, it provides a structured pathway from a general commitment to FAIR towards a progressively harmonized and reusable research data space.
 +*/
  
  
-How to achieve a FAIR state of our data ecosystems+/* How to achieve a FAIR state of our data ecosystems
 What does FAIR mean in that respect? What does FAIR mean in that respect?
  
Line 34: Line 122:
 How do we move forward, what are the priorities? How do we move forward, what are the priorities?
  
-A brief description of the roadmap. +A brief description of the roadmap. */
  
wiki/1_road-to-fair-strategy/start.1785974726.txt.gz · Last modified: by esoeding