Skip to main content
Have a personal or library account? Click to login
Designing Repository-Level Evaluation Criteria: A Data Stewardship Framework for Research Data Repositories Cover

Designing Repository-Level Evaluation Criteria: A Data Stewardship Framework for Research Data Repositories

Open Access
|Aug 2026

Full Article

1. Introduction

Effective data management has become a fundamental requirement in data-intensive research environments, as it underpins the reliability, reusability, and long-term value of research data. The rise of open science and the shift toward a data-intensive research paradigm has highlighted the need for advanced data management strategies. In response, numerous countries are implementing a national strategy for research data sharing and utilization and establishing governance systems to support these policies effectively. Consequently, the role of data repositories is becoming increasingly crucial in meeting diverse requirements for data accessibility, quality, and reliability (Kim, Yang and Kim, 2022; Kim, 2023).

Data repositories are established and operated by national agencies, research institutions, international organizations, disciplinary communities, and commercial providers to store, curate, and share research data. Additionally, to verify data submitted to academic journals, publishers generally require registration in these data repositories. Such data repositories can ultimately become tools that promote data reuse and safely preserve it. The sustainability of data sharing and reuse depends not only on dataset reliability but also on the institutional stewardship mechanisms that govern repository operations. However, it is impossible to review the reliability of all data in a data repository, and therefore, primary reliability can be secured with rich metadata from the data. Another way to ensure data reliability beyond metadata is to secure the quality of the data repository itself. While robust repository governance does not inherently guarantee the intrinsic scientific validity of individual datasets, it establishes institutional and procedural conditions that support authenticity, integrity, traceability, and responsible reuse (Kindling and Strecker, 2022; Han and Han, 2022; Kim, Kang and Kim, 2023).

As of 2024, there are 3,207 data repositories globally, which are highly distributed and have diverse characteristics across domains, regions, organizations, and countries (re3data.org, 2024). This diversity poses significant challenges to data management and sharing, emphasizing the need to effectively evaluate and manage data repositories. Despite the important role data repositories play, there is a lack of research focusing on their evaluation and systematic management.

Efforts by various organizations, such as the FAIRsharing community, Science Europe, and the Confederation of Open Access Repositories (COAR) Working Group, have sought to establish frameworks and standards to improve the reliability and efficiency of data repositories. The FAIRsharing community has presented standards at both the repository and dataset levels. Science Europe has published practical guidelines for the international alignment of research data management and presented criteria for selecting reliable repositories (Science Europe, 2021; Sansone et al., 2020). The COAR Working Group has developed the COAR Community Framework for best practices in repositories. It presents a structured approach to repositories by dividing the characteristics they must have into essential and desired characteristics (COAR, 2022).

The functional requirements for the research data repository platforms are also being studied in various ways. Still, it has been confirmed that the level of requirements applied to the repositories varies depending on the research perspective. The Research Data Alliance (RDA) (Repository Platforms for Research Data IG, 2016) proposed 13 requirements, and Kim (2018) proposed 75 requirements across 13 categories using Purdue University’s data curation profile. Similarly, Webb and McGoohan (2015) described 71 requirements derived from building a digital repository in Ireland. While these studies provide detailed functional requirements for data repositories, there remains a need to further address governance frameworks, metadata standardization, long-term preservation strategies, transparency of repository operations, and quality assurance mechanisms in order to establish more integrated and universally applicable evaluation criteria. To derive additional functional requirements of data repositories, Kim (2020) analyzed Data Management Plans (DMPs) and the CoreTrustSeal certification framework to identify repository-related evaluation criteria. However, while these criteria emphasize compliance, documentation, and trustworthiness standards, they do not fully address cross-institutional governance integration, life cycle-stage differentiation, and structural interoperability across heterogeneous repository environments. Lin, Crabtree and Dillo (2020) provided guidelines to demonstrate the transparency, responsibility, user-centeredness, sustainability, and technology of data repositories. Wu et al. (2019) identified 10 recommendations to improve data searchability and users’ data search experience through research.

Although frameworks suitable for data repositories in specific fields are being developed, they are often limited to their respective domains and thus limited in their ability to describe general datasets. For instance, Liaw et al. (2021) developed a data quality assessment framework across the data life cycle by analyzing 120 research articles. Tilki et al. (2020) defined 19 criteria based on individual clinical trial data. While these studies are instrumental in their respective domains, they are confined to specific domains of data repositories, highlighting the need for more universally applicable frameworks that can handle broader datasets.

In addition to the operational and community-driven frameworks discussed above, a substantial body of work has emerged from certification and trust-oriented initiatives grounded in the Open Archival Information System (OAIS) reference model (ISO 14721). ISO 16363 and its precursor Trustworthy Repositories Audit & Certification: Criteria and Checklist (TRAC) provide formal audit and certification criteria for trustworthy digital repositories. CoreTrustSeal has become the most widely adopted community-based certification mechanism, focusing on organizational infrastructure, digital object management, and technical infrastructure. Similarly, the Nestor seal in Germany and related frameworks emphasize structured assessment procedures for repository trustworthiness. Complementary efforts, such as the Transparency, Responsibility, User focus, Sustainability, and Technology (TRUST) Principles and Preservation Metadata: Implementation Strategies (PREMIS), further articulate governance, transparency, and preservation metadata requirements.

While these canonical frameworks provide robust mechanisms for certification, audit, and governance validation, they are primarily designed as compliance-oriented or principle-based reference systems. The present study does not seek to replicate, replace, or compete with these established certification processes. Rather, it consolidates operational and functional evaluation elements across repository-related sources and introduces an expert-informed prioritization structure that may support routine self-assessment, comparative benchmarking, and preparatory diagnostics prior to formal certification. Table 1 summarizes the positioning of the proposed framework in relation to these established certification and trust-oriented approaches.

Table 1

Positioning of the proposed framework in relation to established certification and trust frameworks.

FRAMEWORKSCOPE AND ROLEASSESSMENT MODEPRIMARY OUTPUTCOVERAGE MAPPINGRELATION TO THIS STUDY
OAIS Reference ModelConceptual reference model for long-term digital preservationConceptual model (non-audit)Preservation frameworkGovernance, preservation processesProvides conceptual foundation for repository trust but does not specify operational evaluation elements
ISO 16363Standard for audit and certification of trustworthy digital repositoriesFormal audit-based certificationCertification decisionOrganizational infrastructure, digital object management, technical infrastructureCertification-focused; does not prioritize operational criteria for routine assessment
TRACChecklist-based audit framework for digital repositoriesAudit checklistTrust evaluation reportOrganizational, technical, preservation controlsPrecursor to ISO 16363; audit-oriented rather than element consolidation
CoreTrustSealCommunity-based repository certification mechanismCertification against defined requirements (R0–R16)Certification sealOrganizational, digital object management, technical infrastructureFocuses on compliance verification rather than cross-source operational integration
TRUST PrinciplesHigh-level governance principles for digital repositoriesPrinciple-based guidanceGovernance principlesTransparency, responsibility, sustainability, user focus, technologyConceptual guidance without prioritized operational breakdown
PREMISPreservation metadata specificationMetadata standardSemantic units for preservation metadataPreservation metadata, provenanceSupports preservation documentation but not holistic repository evaluation
This StudyIntegrated operational evaluation frameworkExpert-validated prioritizationStructured criteria library (Mandatory/Optional)Governance, system support, workflows, QA/QC mechanisms, access and identifiersConsolidates functional elements across sources and introduces prioritization for routine and comparative assessment

This study aims to evaluate repository-level stewardship structures, governance mechanisms, and institutional quality control processes by synthesizing limitations identified in previous frameworks and validating comprehensive evaluation criteria through expert assessment. This framework does not evaluate intrinsic dataset-level scientific quality (e.g., methodological rigor, statistical validity, or completeness). Instead, it assesses repository-level stewardship mechanisms that enable scalable and sustainable quality assurance. By critically analyzing the various requirements presented in existing frameworks and previous studies, this study aims to not only accommodate multiple requirements in the general fields but also set criteria for data repository evaluation. The representative research questions of this study are as follows:

What criteria comprehensively evaluate a data repository?

How relevant and prioritized are these criteria in assessing relevance and priority?

Assessing the relevance and priority of these criteria will help bridge the gap in evaluating data repositories and contribute to a more consistent framework for their management.

2. Research Methodology

In this study, we employed a systematic approach to create and validate criteria for evaluating data repositories. This involved four stages to develop assessment criteria.

Stage 1: Formulation of research question

We initially formulated research questions focused on determining the appropriateness and priorities of various evaluation criteria for the data repository. We investigated existing research data repository evaluation standards and identified limitations.

Stage 2: Case study and analysis

We selected five representative repository-related frameworks and community initiatives, including the RDA, the COAR Community Framework, the Science Europe repository selection criteria, the FAIRsharing registry criteria, and DSpace-based evaluation guidelines. We aimed to identify and document existing evaluation criteria by conducting detailed reviews of their documentation, operational procedures, and published criteria to derive a set of representative evaluation criteria for various repository functionalities and challenges in the next stage.

Stage 3: Empirical investigation with expert validation

This stage emphasized experts conducting an empirical investigation and validation. A survey was conducted with several experts to determine the appropriateness and priority of the derived criteria. The evaluation criteria used Cronbach’s alpha coefficient to ensure consistent internal reliability. Additionally, through priority analysis, based on their average ratings, experts classified elements as essential or optional, providing a balanced framework for repository evaluation.

Stage 4: Framework development through pilot evaluation

We developed a comprehensive framework for data repository evaluation based on the validated criteria in the final stage. The proposed evaluation criteria aim to guide pilot evaluations of data repositories by identifying areas for improvement and additional considerations. We expect this framework to significantly contribute to developing more consistent and effective methods for evaluating and managing data repositories.

3. Case Study and Analysis

To evaluate data storage standards, we applied commonly used evaluation criteria. There are a total of five cases used in this study: functional requirements for RDA, COAR Community Framework for best practices in repositories, repository selection criteria by Science Europe, repository selection and registry criteria by FAIRsharing, and DSpace-based repository evaluation. Since the original text of the evaluation criteria for each case is very large, only the main requirements are summarized here. Full requirements can be found at this link: https://doi.org/10.22747/paper_data.20240711.6.

3.1 Functional requirements of RDA

The Repository Platforms for Research Data Interest Group includes 13 categories and 44 functional requirements. The main requirements of each category are given in Table 2.

Table 2

Main requirements of RDA Data Repository Criteria.

CATEGORYREQUIREMENT
MetadataSupport for various metadata schemas, including domain-specificity and interoperability, enabling data annotation by owners, authorized individuals, or automatic tools, and evaluating metadata quality.
Persistent Identifiers (PID)Assignment of PID/DOI and integration of PIDs into data management.
AuthenticationEnable fine-grained authentication and authorization, allow integration or import from external systems, and provide single sign-on with support for various authentication methods.
Data AccessAllow data providers to choose the access level (e.g., Open Access), offer state-of-the-art interfaces and clients for the repository’s lifetime, provide authorized users access to data versions, enable embargo date selection, offer sophisticated search capabilities, and allow local downloads of selected information.
Policy SupportEnable the automated use of data policies with enforcement points and require all data to be attributed with handling requirements.
PublicationProvide data access statistics through external analytics services or internal monitoring, display bibliographic citations for data with export options for citation software, and maintain citations linked to the data.
Submission/Ingest/ManagementProvide application programming interfaces (APIs) for automated task execution, record audit trails, offer a user-friendly ingest process, ensure integrity and quality control for data and metadata, support micro-services, enable fast data transfer and remote access management, facilitate data and metadata collection with mobile devices, offer definable workflows and both single and batch ingest paths, allow product updates and content deletion by authorized users, provide a vocabulary service, and support staged content.
Data OrganizationResource naming through collection virtualization/logical naming
LocationTight integration with (near) data processing.
IntegrationSupport federation, including storage drivers.
Preservation and SustainabilityMaintain a permanent history of all data versions, convert files to accessible formats while allowing proprietary and legacy types case-by-case, ensure repository scalability, and support workflows.
User Experience/User InterfaceEnsure seamless integration of data and research outputs into a coherent discovery and access solution and allow the creation of exceptional collection views or digital exhibitions.
Data and Product QualityCapture ‘degree of confidence’ on each data item.

3.2 COAR Community Framework for best practices in repositories

The COAR Community Framework provides a global, multidimensional framework that is applicable to various repository types and contexts, divided into essential and desired characteristics. The categories are given in Table 3.

Table 3

Main requirements of COAR Community Framework.

OBJECTIVEESSENTIAL CHARACTERISTICS
DiscoverabilityThe repository supports basic and advanced Dublin Core metadata, Open Archives Initiative Protocol for Metadata Harvesting (OAI-PMH) harvesting, tombstone pages for withdrawn resources, PIDs, search functionality, indexing by external services, inclusion in repository registries, and provides metadata in both human-readable and machine-readable formats.
AccessThe repository offers free access to resources including resource links on landing pages, supports accessibility for persons with disabilities, and provides mechanisms to restrict access to sensitive data for authorized users only.
ReuseThe repository includes licensing information in the metadata record, stipulating reuse conditions for the resource.
Integrity and AuthenticityThe repository employs security practices to prevent unauthorized manipulation, supports metadata revision and resource versioning, and regularly performs integrity checks to detect unauthorized changes or accidental damage.
Quality AssuranceThe repository conducts lightweight reviews and enhancements of basic metadata upon submission and provides documentation or policies detailing the applied curation processes for resources and metadata.
PreservationThe repository has a digital preservation plan detailing management duration, roles, and procedures, records checksums upon submission or modification, collects preservation metadata, ensures that depositor agreements cover necessary preservation actions, allows metadata and resources to be copied or migrated, stores copies in different locations, and maintains a business continuity plan for natural disasters or cyber-attacks.
Sustainability and GovernanceThe repository identifies its managing organization and governance, provides user assistance and dedicated management staff, responds to queries promptly, has a policy for resource management if operations cease, and maintains a long-term management and funding plan.
OtherThe repository provides public documentation that outlines the scope of the resources accepted in the repository.

3.3 Repository selection criteria by Science Europe

The repository selection criteria by Science Europe emphasizes selecting reliable repositories organized into four categories (Table 4).

Table 4

Main requirements of Science Europe data repository criteria.

CATEGORYREQUIREMENT
Provision of PIDsAllow data discovery and identification, enable searching, citing, and retrieval, and provide support for data versioning.
MetadataEnable finding, referencing, and providing publicly available data, including protected or retracted data, using broadly accepted metadata standards that are machine-retrievable.
Data Access and Usage LicensesEnable access to data under well-specified conditions, ensure data authenticity and integrity, enable data retrieval, provide licensing and permissions information (ideally in machine-readable form), and ensure confidentiality and rights of data subjects and creators.
PreservationEnsure persistence of metadata and data and be transparent about mission, scope, preservation policies, and plans, including governance, financial sustainability, retention period, and continuity plan.

3.4 Repository selection criteria by FAIRsharing community

The FAIRsharing Initiative presents repository-level and dataset-level criteria included in Table 5.

Table 5

Main requirements of FAIRsharing Initiative.

CRITERIADEFINITION
Certification and Community BadgesDoes the repository have certification schemes or community badges assessing its fitness, trustworthiness, and adoption?
Data Access ConditionsWhat is the process for requesting and granting access to the repository or dataset, and are the data freely available or subject to a request and approval process?
Data and Metadata Standards What community-defined standards does the repository implement to ensure consistent, machine-readable representation of data and/or metadata, facilitating their discovery and interpretation?
Data CurationAre there minimum curation steps that the repository performs on submitted data, and is there a webpage or document describing the type of curation done?
FundingWhat type of funding (e.g., grants, donations, memberships) does the repository receive, and which organizations fund it?
PIDs for DataDoes the repository assign globally unique and persistent identifiers to the deposited data, and what identifier schema is used?
Repository CoverageWhat higher-level subject areas, disciplines, and cross-disciplinary domains does the repository cover, including data types, technology, and study?
Repository StatusWhat is the life cycle status of the repository: Is it still being developed, or is it in production and accepting data submissions while possibly undergoing (re)development, enhancement, or maintenance?
Resource SustainabilityDoes the repository have a webpage or document that describes its sustainability plans?
User SupportDoes the repository provide a contact point, such as a helpdesk email or contact form, to assist data depositors and users during or after submission?
FAIRsharing Record MaintainerHas the owner or maintainer of the repository claimed the record in FAIRsharing and vetted its descriptions, ensuring verified information and tracking the resource’s evolution, alongside FAIRsharing’s in-house curation?

3.5 DSpace-based repository platform evaluation guidelines

Tilki et al. (2020) proposed 19 quality criteria for clinical data repositories, which were categorized into the following in Table 6.

Table 6

Main requirements of DSpace-based repository evaluation.

REQUIREMENTDESCRIPTION
Guidelines for Data Upload and StorageSupporting various file types and metadata schemas and providing upload mechanisms with instructions.
De-identification Practices Before UploadProviding links to de-identification tools and requirements and implementing de-identification tools.
Control of Data QualitySupporting quality control in its workflow.
Formal Contract for Upload and StorageIncorporating a data transfer agreement in the system workflow.
Application of a Metadata SchemaUsing a consistent metadata schema, allowing customization, providing tools for metadata completion, and making metadata publicly available.
Application of an IdentifierApplying a primary PID system and using other PIDs as appropriate.
Flexibility of AccessAllowing open access with optional embargo, web-based self-attestation, managed access through group membership or case-by-case basis, and supporting granular access to different dataset parts.
Long-Term PreservationSupporting long-term data and metadata preservation using sustainable software systems.

3.6 Synthesis of repository evaluation criteria

We analyzed five existing cases to design comprehensive evaluation criteria for evaluating data repositories. The RDA presents the most comprehensive criteria, encompassing 13 categories, and addresses almost all aspects except certification, community badges, funding, repository coverage, and record maintainers. The COAR prioritizes discoverability, access, reuse, and governance but lacks specific criteria for metadata, authentication, policy support, publication, location, integration, and de-identification. Science Europe emphasizes metadata, PIDs, data access, and preservation but does not cover several other critical areas, such as policy support, publication, and quality assurance. The FAIRsharing community includes a wide range of criteria, featuring unique elements like certification, community badges, and funding, yet it needs detailed criteria for submission, integration, and location. The DSpace-based repository platform evaluation guidelines have specific criteria for de-identification, flexibility of access, and formal contracts, but they need to include broader criteria found in other frameworks.

To address these discrepancies, we proposed comprehensive evaluation criteria integrating the categories and requirements emphasized by these five cases, mainly based on the standard and widely utilized criteria of the RDA (Table 7). A total of 13 requirements and 150 elements were identified among the five cases. Specifically, RDA originally had 44 elements, but due to overlapping elements, we expressed it as 46. Similarly, COAR consisted of 52 elements, but 56 were used in this study for the same reason as RDA.

Table 7

Distribution of evaluation criteria elements across five cases.

CATEGORYREQUIREMENTNUMBER OF ELEMENTS
RDACOARSCIENCE EUROPEFAIRSHARINGDSPACETOTAL
A Data Access and RetrievalA-1Metadata3351416
A-2Access and Retrieval62031535
B Permanent IdentifierB-1Permanent Identifier223119
C Sustainability and GovernanceC-1Rights, License and Copyright1225
C-2Preservation and Sustainability41321121
C-3Policy Support and Repository Coverage334111
C-4Repository Certification11
C-5Funding112
D Data Curation and CitationD-1Curation of Data15116
D-2Data Citation314
E User Interface and System SupportE-1User Interface426
E-2System Support342615
F Quality ControlF-1Data Authenticity and Integrity25119

The distribution of requirements covered by each case is as follows: RDA and COAR covered 11 requirements each, FAIRsharing covered eight, and both Science Europe and DSpace covered seven requirements each. Notably, the category A-2 Access and Retrieval had the highest number of requirements, with the five cases covering 35 requirements. In contrast, the category C-5 Funding had the fewest requirements, with only two.

4. Empirical Investigation with Expert Validation

The purpose of the survey was to establish evaluation criteria for data repositories and determine the appropriateness and priorities of these criteria. The questionnaire was distributed by e-mail in May 2023 to 40 experts with demonstrated experience in research data repositories, research data management, metadata, and data curation. The participants were identified through existing professional collaboration networks established through previous research activities. They represented universities, government-funded research institutes, private companies, and non-profit organizations involved in repository development, operation, and research data management. Participation was voluntary and all responses were collected anonymously. No personally identifiable information or sensitive personal information was collected. A total of 32 experts responded, resulting in an 80% response rate. The respondents were affiliated with various institutions: 17 from research institutes, 10 from universities, four from private companies, and one from a non-profit organization. Regarding research experience, 16 had more than 20 years, 10 had 10–20 years, four had less than 5 years, and two had 5–10 years. The respondents’ fields included 17 in science and engineering, 11 in humanities and social sciences, three in science, and one who did not respond.

The questionnaire consisted of six categories rated on a 7-point Likert scale: Data Access and Retrieval, Permanent Identifier, Sustainability and Governance, Data Curation and Citation, User Interface and System Support, and Quality Control.

The validity of the response structure was verified by obtaining an index of reliability. Cronbach’s alpha was used to measure the consistency and homogeneity of the items, with values between 0 and 1, where being closer to 1 indicates higher reliability. If k is the number of items in the target item, σi2 is the variance of each item, and σt2 is the variance of the item to which the item belongs, the Cronbach’s alpha coefficient can be calculated as kk1(1σi2σt2). The Cronbach’s alpha coefficients for each category are as shown in Table 8.

Table 8

Cronbach’s alpha coefficients for categories.

CATEGORYALPHA COEFFICIENT
A Data Access and Retrieval0.82
B Permanent Identifier0.76
C Sustainability and Governance0.94
D Data Curation and Citation0.90
E User Interface and System Support0.93
F Quality Control0.76

The results showed that the Cronbach’s alpha coefficient for all categories was over 0.7, indicating that the responses were well-structured and consistent, thus reliable. The survey results contained the importance of each evaluation factor for the data repository evaluation criteria by researchers. Given the high reliability and validity of the survey responses, it was deemed reasonable to establish the data repository evaluation criteria based on these results.

To address the challenge of reflecting the opinions of all researchers in the field, a resampling approach was employed due to the limited number of respondents. By obtaining the sample mean from the survey responses and applying replacement sampling repeatedly, the distribution of the sample mean can be approximated to a normal distribution according to the Central Limit Theorem. The average of these sample means can be estimated as the average response of the entire population of researchers in the related field.

5. Framework Development through Pilot Evaluation

We verified the response averages and priorities of each element of requirements for the comprehensive evaluation criteria for the data repository through empirical investigation. Given that we consistently derive the response average and priority of each element of requirements across all categories, we focus here on the ‘C Sustainability and Governance’ category, which had the highest Cronbach’s alpha coefficients. Detailed information on the response averages and priorities of each element in other categories can be found in the Supplementary file (https://doi.org/10.22747/paper_data.20240711.6).

For each requirement in the ‘C Sustainability and Governance’ category, the priority of the sub-elements was analyzed (Figure 1). The response average and results for each element were graphically displayed, and those with scores lower than the average were designated as optional, while those with scores higher than the average were designated as mandatory (Table 9).

Figure 1

Response score for elements of ‘C Sustainability and Governance’ category.

Table 9

Response average and priority for elements of ‘C Sustainability and Governance’ category.

REQUIREMENTELEMENTAVERAGE SCOREOBLIGATION
C-1
Rights, License and Copyright
1Provide information about licensing and permissions ideally in machine-readable form.6.38Mandatory
2Ensure confidentiality and rights of data subjects and creators.6.28Mandatory
3Allow open access to material, with an optional embargo period.5.97Mandatory
4Offer managed access through group membership.5.16Optional
5Support granular access to different parts of dataset collections.4.63Optional
6The resources in the repository are available at no cost to the user.5.34Optional
7Provide different access rights for groups and individuals (roles) on collections and allow the import of such concepts (e.g., from identity management systems). In the case of confidential or proprietary data, authenticate every access and authorize every operation.5.53Optional
8Provide single sign-on and/or support for different authentication methods.5.22Optional
C-2
Preservation and Sustainability
1Support long-term preservation of data and metadata.6.28Mandatory
2The repository collects basic preservation metadata including provenance, date of upload, and file format.5.63Mandatory
3The metadata and the resources in the repository can be copied or migrated to other systems.5.47Optional
4At least one copy of the repository contents is stored in a different location than the original repository.5.66Mandatory
5The agreement between depositor and repository provides for all actions necessary to meet preservation responsibilities—for example, rights to copy, transform, and store the items.5.78Mandatory
6Ensure persistence of metadata and data.6.41Mandatory
7Files need to be converted to the most accessible formats.6.28Mandatory
C-3
Policy Support and Repository Coverage
1Data policies are used to define what happens when to which dataset. For example, for processing and quality control, regularly enforced policies are helpful.5.94Mandatory
2Policy enforcement points: Control all operations with administrator-defined rules.5.68Mandatory
3The higher-level subject areas/disciplines that the repository covers, as well as cross-disciplinary domains, such as the types of data, technology, and study.4.94Optional
4The repository provides public documentation that outlines the scope of the resources accepted in the repository.5.16Optional
5The life cycle status of the repository: Is it still being developed or is it in production and accepting data submissions? The latter does not exclude that some (re)development, enhancement, or maintenance may be ongoing, as happens in any repository.5.19Optional
6Be transparent about mission, scope, preservation policies, and plans (including governance, financial sustainability, retention period, and continuity plan).5.97Mandatory
7The repository has a digital preservation plan that states the duration of time that the resources will be managed for, identifies roles, and documents procedures for the preservation of different resource formats.5.72Mandatory
8The repository has a business continuity plan that details the response and procedures in case of natural disasters or cyber-attacks.6.41Mandatory
9Plan that gives information about sustainability plans for the repository: Does the repository have a webpage or document that describes these?5.34Optional
10The repository provides documentation or has a policy outlining what curation processes are applied to the resources and the metadata.5.35Optional
11The repository is included in one or more disciplinary or general registry of repositories.5.48Optional
C-4
Repository Certification
1Certification schemes and/or community badges that assess certain aspects of the repository (e.g., its fitness, trustworthiness, adoption): Does the repository have any?5.31Optional
C-5
Funding
1The type of funding (e.g., grants, donations, memberships) and the organization(s) that fund the repository.5.00Optional
2The repository (or organization that manages the repository) has a long-term plan for managing and funding the repository.5.03Optional

For the C-1 Rights, License, and Copyright requirements, the averages of elements C-1-1, C-1-2, and C-1-3 were 6.38, 6.28, and 5.97, respectively. These scores were higher than the average for the entire C category; thus, we designated these elements as mandatory items. Notably, these elements had an average standard deviation of 1.03, indicating relatively small variability among respondents. Elements C-1-4 and C-1-5 scored below average, so we designated these elements as optional items. For the C-2 requirements (Preservation and Sustainability), most elements scored above average, indicating strong sustainability. For the C-3 requirements (Policy Support and Repository Coverage), there was a mix of high and low scores, suggesting variability in policy support and repository coverage. Elements C-3-3 and C-3-4 scored below average and were designated as optional, while C-3-1, due to its high score, was designated as a mandatory item. The C-4 requirement (Repository Certification) had a score for element C-4-1 that was below average, with a standard deviation of 1.09, indicating low variability and consensus among most respondents. The two elements of the C-5 requirement (Funding), C-5-1 and C-5-2, exhibited high variability and scored below average; thus, both were designated optional items.

6. Proposal of Comprehensive Evaluation Criteria for Data Repositories

The proposed criteria emphasize repository-level stewardship capacity, including governance transparency, preservation planning, integrity safeguards, provenance tracking, workflow control, and sustainable infrastructure. The framework focuses on structural enablers of quality assurance rather than direct scientific validation of dataset content. The comprehensive evaluation criteria comprise six categories, 13 requirements, and 110 elements, including mandatory and optional obligations (Table 10). The ‘E User Interface and System Support’ category had the most elements, totaling 38, while the ‘B Permanent Identifier’ category had the fewest, with only six. This number does not indicate the importance of the categories but rather the number of independent requirements within each category.

Table 10

Comprehensive evaluation criteria for data repositories.

CATEGORYREQUIREMENTSELEMENTSSOURCEOBLIGATION
A Data Access and RetrievalA-1
Metadata
1Support various types of metadata.RDAMandatory
2Ensure that metadata are machine-retrievable.Science Europe
3Provide information that is publicly available and maintained, even for non-published, protected, retracted, or deleted data.Science EuropeOptional
A-2
Access and Retrieval
1Include a link to each resource on its respective landing page in the repository.COARMandatory
2Enable linking between related contents in the metadata record, such as preprints, published articles, data, and software.COAR
3Allow data providers to choose the access level for data (e.g., Open Access).RDA
4Enable data or at least metadata retrieval using an open and standardized protocol.Science Europe
5Permit local download of selected information.RDA
6Support the use of controlled vocabularies in metadata records.COAROptional
7Facilitate indirect access to restricted resources (e.g., by contacting the author).COAR
8Allow data-depositing users to select an embargo date.RDA
9Provide a tombstone page for withdrawn resources, ensuring the metadata record remains publicly available.COAR
10Offer support to users during or after submission with a contact point (e.g., helpdesk email or contact form) to assist data depositors and users.FAIR sharing
B Permanent IdentifierB-1
Permanent Identifier
1Assign a PID/DOI to data and collections during data ingestion or ‘project publication’ time or earlier (e.g., when a paper is submitted but the data are not final yet). Ensure that the PID resolves the research data’s ‘landing page,’ displaying the required descriptive metadata during the embargo period, and establish a clear transition from PID collections to a DOI.RDAMandatory
2Integrate all data management activities with PID management. Ensure that PID metadata is always synchronized with the data/metadata holdings.RDA
3Ensure that PIDs are included in the corresponding metadata.Science Europe
4Consistently assign PIDs (e.g., DOI, URN, ARK) to the data, allowing the corresponding data and metadata to be found, referred to, and retrieved, even if the storage location changes.Science Europe
5Support PIDs for authors, funders, institutions, funding programs, and other relevant entities.COAROptional
6Ensure that metadata information can declare links to other relevant or associated information by providing the PID and a description of the scientific relation, including details of the associated researcher, with permanent research IDs (e.g., ORCID, ISNI, DAI).Science Europe
C
Sustainability and Governance
C-1
Rights, License and Copyright
1Provide information about licensing and permissions, ideally in machine-readable form.Science EuropeMandatory
2Ensure confidentiality and rights of data subjects and creators.Science Europe
3Allow open access to material, with an optional embargo period.DSpace
4Offer managed access through group membership.DSpaceOptional
5Support granular access to different parts of dataset collections.DSpace
6Ensure that resources in the repository are available at no cost to the user.COAR
7Provide different access rights for groups and individuals (roles) on collections, allowing the import of such concepts (e.g., from identity management systems). Authenticate every access and authorize every operation for confidential or proprietary data.RDA
8Provide single sign-on and/or support for different authentication methods.RDA
C-2
Preservation and Sustainability
1Support long-term preservation of data and metadata.DSpaceMandatory
2Collect basic preservation metadata, including provenance, date of upload, and file format.COAR
3Store at least one copy of the repository contents in a different location than the original repository.COAR
4Ensure that the agreement between depositor and repository provides for all actions necessary to meet preservation responsibilities (e.g., rights to copy, transform, and store the items).COAR
5Ensure persistence of metadata and data.Science Europe
6Convert files to the most accessible formats.RDA
7Ensure that metadata and resources in the repository can be copied or migrated to other systems.COAROptional
C-3
Policy Support and Repository Coverage
1Use data policies to define what happens to which dataset, enforcing policies regularly for processing and quality control.RDAMandatory
2Implement policy enforcement points to control all operations with administrator-defined rules.RDA
3Be transparent about mission, scope, preservation policies, and plans (including governance, financial sustainability, retention period, and continuity plan).Science Europe
4Have a digital preservation plan that states the duration of time that the resources will be managed for, identifies roles, and documents procedures for the preservation of different resource formats.COAR
5Maintain a business continuity plan detailing the response for various scenarios.COAR
6Specify the higher-level subject areas/disciplines that the repository covers, including cross-disciplinary domains, types of data, technology, and study.FAIR sharingOptional
7Provide public documentation outlining the scope of the resources accepted in the repository.COAR
8Indicate the life cycle status of the repository: whether it is still being developed or is in production and accepting data submissions.FAIR sharing
9Provide information about sustainability plans for the repository, such as a webpage or document describing these.FAIR sharing
10Document or have a policy outlining the curation processes applied to resources and metadata.COAR
11Ensure that the repository is included in one or more disciplinary or general registry of repositories.COAR
C-4
Repository Certification
1Obtain certification schemes and/or community badges that assess certain aspects of the repository (e.g., fitness, trustworthiness, adoption).FAIR sharingOptional
C-5
Funding
1Specify the type of funding (e.g., grants, donations, memberships) and the organization(s) that fund the repository.FAIR sharingOptional
2Ensure that the repository (or managing organization) has a long-term plan for managing and funding the repository.FAIR sharing
D
Data Curation and Citation
D-1
Curation of Data
1Maintain a permanent history of versions for all data.RDAMandatory
2Ensure that the version of the data stored in the repository is clearly specified and documented via a permanent audit trail for provenance tracing.Science Europe
3Support the revision of metadata and versioning of resources.COAR
4Incorporate a formal contract regarding upload and storage, including a data transfer agreement in the system workflow.DSpace
5Handle staged content, including submission states that are raw, processed, curated, and published.RDAOptional
6Provide an easy-to-use ingest process with minimal barriers to participation.RDA
7Define a submission/ingest workflow.RDA
8Perform review and annotation of data, ensuring that a set of minimum curation steps are applied to the submitted data.FAIR sharing
9Provide a webpage or document describing the type of curation done.FAIR sharing
10Register workflows as executable objects and track the provenance of each workflow execution.RDA
11Define data access mechanisms and terms at the repository and/or dataset level, outlining the process for requesting and granting access.FAIR sharing
12Provide different versions of a dataset.RDA
D-2
Data Citation
1Display bibliographic citations for data and allow exporting bibliographic data to citation software (e.g., EndNote, Citavi, Zotero).RDAMandatory
2Ensure that citations provide recognition and updates from others utilizing the data for experiments or other purposes.RDA
3Make metadata in the repository available for download in a standard bibliographic format at no cost to the user.COAROptional
E
User Interface and System Support
E-1
User Interface
1Support a responsive, mobile-friendly user interface.COARMandatory
2Provide interfaces (APIs) for automated task execution, such as data ingestion or integration with data analysis tools and other external applications.RDA
3Ensure that scientific terms are consistent for future reuse via a vocabulary service.RDA
4Provide data access statistics using external analytics services or internal monitoring of user activity.RDA
5Offer sophisticated search capabilities for metadata and data, including full-text search and schema-specific search for both humans and computers.RDA
6Ensure access to documentation and metadata for individuals.COAR
7Allow the creation of special collection views or digital exhibitions.RDAOptional
8Enable fast data transfer, ingestion, and export.RDA
9Support data and metadata collection with mobile devices.RDA
10Allow authorized users to mark content for deletion.RDA
11Collect and share usage information using a standard methodology (e.g., number of views, downloads).COAR
12Apply a metadata schema to describe contents and provide tools to help data generators complete metadata fields.DSpace
13Allow data annotation by the data owner, authorized individuals, or automatic metadata extraction tools.RDA
E-2
System Support
1Support metadata harvesting using OAI-PMH.COARMandatory
2Implement de-identification practices before data upload.DSpace
3Ensure that mechanisms are in place to limit access to authorized users only for sensitive research data.COAR
4Recommend tools to anonymize sensitive data to enable data sharing.COAR
5Provide mechanisms to make very large files available to users outside the normal user interface when file size becomes unwieldy.COAR
6Adhere to the most recent version of the W3C Web Content Accessibility Guidelines.COAR
7Ensure that the repository is scalable regarding the amount of data.RDA
8Use sustainable software systems.DSpace
9Record audit trails to track changes to resource metadata and information relationships.RDA
10Integrate data and other research outputs into a coherent and consistent discovery and access solution.RDA
11Support both individual and bulk uploads in the repository’s submission system.COAR
12Apply a primary PID system.DSpace
13Require all data to be attributed with handling requirements, including licenses and security parameters.RDA
14Support a range of file types and metadata schema for data upload and storage.DSpace
15Provide state-of-the-art user interfaces and clients throughout the repository platform’s lifetime, updating features and functionality to meet current researchers’ requirements and expectations.RDAOptional
16Allow open access after web-based self-attestation of the user.DSpace
17Offer managed access through application on a case-by-case basis.DSpace
18Encapsulate operations in micro-services that can be chained into a workflow.RDA
19Provide both single and batch ingest paths to efficiently submit a range of data types and scales.RDA
20Allow product developers to update product information within the repository.RDA
21Manage data collections and their properties independently of the storage system and resource naming through collection virtualization/logical naming.RDA
22Integrate with near-data processing facilities like High Performance Computing to handle large data volumes efficiently, maintaining provenance information.RDA
23Build the repository on well-supported, open-source software.COAR
24Map access protocol to storage protocol using storage drivers.RDA
25Allow authorized individuals to curate materials from distributed locations through remote access management.RDA
F
Quality Control
F-1
Data Authenticity and Integrity
1Enforce metadata quality evaluation using metrics.RDAMandatory
2Perform regular integrity checks of resources to detect unauthorized changes or accidental damage.COAR
3Ensure data authenticity and integrity by including detailed provenance information in the metadata.Science Europe
4Support quality control in its workflow.DSpace
5Undertake lightweight review and enhancement of basic metadata upon submission of resources.COAROptional
6Record the checksum when a resource is submitted or modified.COAR
7Apply security practices to prevent unauthorized manipulation of resources.COAR
8Capture a structured stewardship confidence indicator derived from curation status, validation stage, provenance completeness, and workflow review history, rather than a direct scientific quality score.RDA
9Maintain a contact person or organization as the FAIRsharing record maintainer for the repository’s description in FAIRsharing, ensuring that the record is claimed and vetted.FAIR sharing

The evaluation criteria consist of 110 elements—55 mandatory and 55 optional. Six of the seven elements of the ‘C-2 Conservation and Sustainability’ requirement are mandatory, making up 86% of the requirements. This means that most factors must be taken into consideration in this category. On the other hand, in the requirements of ‘C-4 Repository Certification’ and ‘C-5 Funding,’ all elements are optional. The focus on mandatory elements helps prioritize essential criteria that data repositories must meet to ensure core functionality and compliance. In contrast, optional elements identify additional features that can enhance repository performance and user satisfaction. Mandatory criteria are non-negotiable for operational standards, whereas optional criteria provide valuable enhancements without being essential for basic functionality.

Among the five cases, the elements created based on RDA were the most numerous, with 42 elements. These elements were adopted by COAR, Dspace, Science Europe, and FAIRsharing in the following order: 32, 13, 13, and 11, respectively. This indicates that RDA includes the broadest categories and the most detailed criteria, while FAIRsharing had fewer elements adopted due to less detailed criteria.

7. Discussion

The findings of this study should be interpreted in relation to the established certification and trust-oriented frameworks for digital repositories. Audit-based standards such as ISO 16363 and CoreTrustSeal provide robust and internationally recognized mechanisms for assessing compliance with predefined requirements and organizational controls. These frameworks play a critical role in ensuring repository trustworthiness and long-term preservation governance.

The framework proposed in this study does not aim to replace or replicate such certification mechanisms. Rather, it integrates functional and governance-related requirements extracted from multiple repository-related sources and introduces an expert-validated prioritization mechanism. In this regard, the framework may serve as an operational assessment instrument supporting routine self-evaluation, comparative benchmarking, and preparatory diagnostics prior to formal certification processes.

The distinction between stewardship quality and intrinsic dataset validity is also conceptually significant. The proposed criteria do not seek to evaluate scientific correctness at the content level; instead, they assess infrastructural and governance conditions that enable transparency, traceability, integrity, and reliability within repository operations. Repository-level mechanisms therefore function as structural enablers of trust rather than substitutes for domain-specific validation procedures.

Finally, the introduction of mandatory and optional classifications differentiates this framework from binary certification models. These classifications reflect empirically derived prioritization patterns and may facilitate incremental improvement strategies across heterogeneous institutional and disciplinary contexts.

The framework may also contribute to ongoing international initiatives aimed at standardizing repository service characteristics, including the RDA Community-based catalogue of requirements for trustworthy Technical Repository Service Providers Working Group (TRSPs WG) and related EU-funded projects such as FIDELIS and EDEN, by providing a structured operational synthesis and prioritization perspective.

8. Conclusion

This study aimed to develop a comprehensive repository-level evaluation framework for assessing data stewardship capacity in research data repositories. Effective data management supports data sharing and reuse, both of which depend on adequate dataset-level quality. Because directly assessing the intrinsic quality of all individual datasets and all quality dimensions (e.g., accuracy, completeness, and consistency) is often impractical at scale, evaluating repository-level stewardship capacity provides a structured and scalable approach to strengthening institutional conditions for trustworthy data management. Such evaluation does not replace intrinsic dataset-level scientific quality assessment but enhances confidence in the governance, preservation, provenance, and integrity mechanisms that underpin sustainable data reuse.

This research analyzed five existing repository evaluation frameworks, formulating 13 requirements across six categories encompassing 110 detailed elements. These experts validated these criteria through a survey, confirming their internal consistency, with all items scoring a Cronbach’s alpha coefficient above 0.7.

The survey results highlighted the importance and prioritization of these criteria, categorizing 55 elements as essential (mandatory) and 55 as additional (optional). The six categories identified include Data ‘Access and Retrieval,’ ‘Permanent Identifier,’ ‘Sustainability and Governance,’ ‘Data Curation and Citation,’ ‘User Interface and System Support,’ and ‘Quality Control.’ Each category was designed to evaluate distinct dimensions of repository-level stewardship rather than intrinsic dataset-level scientific validity.

Despite the limitation of relying on five case studies for criteria derivation and the relatively small sample size of experts (survey participants), this study lays a foundational groundwork for data repository evaluation. The proposed criteria can serve as a preliminary framework, fostering further research and refinement through pilot evaluations. Future studies can build upon these findings to enhance and expand the criteria, which can help ensure comprehensive and effective data repository evaluations. This work aims to contribute significantly to the development of more consistent and effective methods for evaluating and managing data repositories, ultimately supporting better data management practices and the advancement of open science.

Ethics and Consent

This study was based on an anonymous online questionnaire survey conducted in May 2023 to obtain expert opinions for validating the proposed repository evaluation criteria. The participants were recruited through existing professional collaboration networks and consisted of experts with experience in research data repositories, research data management, metadata, and data curation. The survey was conducted solely for research purposes. Participation was voluntary, and respondents were informed of the purpose of the study before completing the questionnaire. Completion of the questionnaire was considered to indicate informed consent to participate in the study. No personally identifiable information or sensitive personal information was collected, and all responses were analyzed in anonymized and aggregated form. According to the institutional policy in effect at the time the study was conducted, a formal Institutional Review Board (IRB) approval was not required for this study.

Data Accessibility Statement

The survey data supporting the findings of this study are available from the corresponding author upon reasonable request.

Author Contributions

Sooyeon Han: Conceptualization, formal analysis, writing – original draft, writing – review and editing.

Jong-Gyu Han: Funding acquisition, writing – review and editing.

Sun-tae Kim: Investigation.

Ju-seop Kim: Investigation, methodology, writing – original draft, writing – review and editing.

All authors have read and agreed to the published version of the manuscript.

Language: English
Page range: 33 - 33
Submitted on: Aug 8, 2025
Accepted on: Jul 31, 2026
Published on: Aug 21, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Sooyeon Han, Suntae Kim, JongGyu Han, Juseop Kim, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.