1. Introduction
Effective data management has become a fundamental requirement in data-intensive research environments, as it underpins the reliability, reusability, and long-term value of research data. The rise of open science and the shift toward a data-intensive research paradigm has highlighted the need for advanced data management strategies. In response, numerous countries are implementing a national strategy for research data sharing and utilization and establishing governance systems to support these policies effectively. Consequently, the role of data repositories is becoming increasingly crucial in meeting diverse requirements for data accessibility, quality, and reliability (Kim, Yang and Kim, 2022; Kim, 2023).
Data repositories are established and operated by national agencies, research institutions, international organizations, disciplinary communities, and commercial providers to store, curate, and share research data. Additionally, to verify data submitted to academic journals, publishers generally require registration in these data repositories. Such data repositories can ultimately become tools that promote data reuse and safely preserve it. The sustainability of data sharing and reuse depends not only on dataset reliability but also on the institutional stewardship mechanisms that govern repository operations. However, it is impossible to review the reliability of all data in a data repository, and therefore, primary reliability can be secured with rich metadata from the data. Another way to ensure data reliability beyond metadata is to secure the quality of the data repository itself. While robust repository governance does not inherently guarantee the intrinsic scientific validity of individual datasets, it establishes institutional and procedural conditions that support authenticity, integrity, traceability, and responsible reuse (Kindling and Strecker, 2022; Han and Han, 2022; Kim, Kang and Kim, 2023).
As of 2024, there are 3,207 data repositories globally, which are highly distributed and have diverse characteristics across domains, regions, organizations, and countries (re3data.org, 2024). This diversity poses significant challenges to data management and sharing, emphasizing the need to effectively evaluate and manage data repositories. Despite the important role data repositories play, there is a lack of research focusing on their evaluation and systematic management.
Efforts by various organizations, such as the FAIRsharing community, Science Europe, and the Confederation of Open Access Repositories (COAR) Working Group, have sought to establish frameworks and standards to improve the reliability and efficiency of data repositories. The FAIRsharing community has presented standards at both the repository and dataset levels. Science Europe has published practical guidelines for the international alignment of research data management and presented criteria for selecting reliable repositories (Science Europe, 2021; Sansone et al., 2020). The COAR Working Group has developed the COAR Community Framework for best practices in repositories. It presents a structured approach to repositories by dividing the characteristics they must have into essential and desired characteristics (COAR, 2022).
The functional requirements for the research data repository platforms are also being studied in various ways. Still, it has been confirmed that the level of requirements applied to the repositories varies depending on the research perspective. The Research Data Alliance (RDA) (Repository Platforms for Research Data IG, 2016) proposed 13 requirements, and Kim (2018) proposed 75 requirements across 13 categories using Purdue University’s data curation profile. Similarly, Webb and McGoohan (2015) described 71 requirements derived from building a digital repository in Ireland. While these studies provide detailed functional requirements for data repositories, there remains a need to further address governance frameworks, metadata standardization, long-term preservation strategies, transparency of repository operations, and quality assurance mechanisms in order to establish more integrated and universally applicable evaluation criteria. To derive additional functional requirements of data repositories, Kim (2020) analyzed Data Management Plans (DMPs) and the CoreTrustSeal certification framework to identify repository-related evaluation criteria. However, while these criteria emphasize compliance, documentation, and trustworthiness standards, they do not fully address cross-institutional governance integration, life cycle-stage differentiation, and structural interoperability across heterogeneous repository environments. Lin, Crabtree and Dillo (2020) provided guidelines to demonstrate the transparency, responsibility, user-centeredness, sustainability, and technology of data repositories. Wu et al. (2019) identified 10 recommendations to improve data searchability and users’ data search experience through research.
Although frameworks suitable for data repositories in specific fields are being developed, they are often limited to their respective domains and thus limited in their ability to describe general datasets. For instance, Liaw et al. (2021) developed a data quality assessment framework across the data life cycle by analyzing 120 research articles. Tilki et al. (2020) defined 19 criteria based on individual clinical trial data. While these studies are instrumental in their respective domains, they are confined to specific domains of data repositories, highlighting the need for more universally applicable frameworks that can handle broader datasets.
In addition to the operational and community-driven frameworks discussed above, a substantial body of work has emerged from certification and trust-oriented initiatives grounded in the Open Archival Information System (OAIS) reference model (ISO 14721). ISO 16363 and its precursor Trustworthy Repositories Audit & Certification: Criteria and Checklist (TRAC) provide formal audit and certification criteria for trustworthy digital repositories. CoreTrustSeal has become the most widely adopted community-based certification mechanism, focusing on organizational infrastructure, digital object management, and technical infrastructure. Similarly, the Nestor seal in Germany and related frameworks emphasize structured assessment procedures for repository trustworthiness. Complementary efforts, such as the Transparency, Responsibility, User focus, Sustainability, and Technology (TRUST) Principles and Preservation Metadata: Implementation Strategies (PREMIS), further articulate governance, transparency, and preservation metadata requirements.
While these canonical frameworks provide robust mechanisms for certification, audit, and governance validation, they are primarily designed as compliance-oriented or principle-based reference systems. The present study does not seek to replicate, replace, or compete with these established certification processes. Rather, it consolidates operational and functional evaluation elements across repository-related sources and introduces an expert-informed prioritization structure that may support routine self-assessment, comparative benchmarking, and preparatory diagnostics prior to formal certification. Table 1 summarizes the positioning of the proposed framework in relation to these established certification and trust-oriented approaches.
Table 1
Positioning of the proposed framework in relation to established certification and trust frameworks.
| FRAMEWORK | SCOPE AND ROLE | ASSESSMENT MODE | PRIMARY OUTPUT | COVERAGE MAPPING | RELATION TO THIS STUDY |
|---|---|---|---|---|---|
| OAIS Reference Model | Conceptual reference model for long-term digital preservation | Conceptual model (non-audit) | Preservation framework | Governance, preservation processes | Provides conceptual foundation for repository trust but does not specify operational evaluation elements |
| ISO 16363 | Standard for audit and certification of trustworthy digital repositories | Formal audit-based certification | Certification decision | Organizational infrastructure, digital object management, technical infrastructure | Certification-focused; does not prioritize operational criteria for routine assessment |
| TRAC | Checklist-based audit framework for digital repositories | Audit checklist | Trust evaluation report | Organizational, technical, preservation controls | Precursor to ISO 16363; audit-oriented rather than element consolidation |
| CoreTrustSeal | Community-based repository certification mechanism | Certification against defined requirements (R0–R16) | Certification seal | Organizational, digital object management, technical infrastructure | Focuses on compliance verification rather than cross-source operational integration |
| TRUST Principles | High-level governance principles for digital repositories | Principle-based guidance | Governance principles | Transparency, responsibility, sustainability, user focus, technology | Conceptual guidance without prioritized operational breakdown |
| PREMIS | Preservation metadata specification | Metadata standard | Semantic units for preservation metadata | Preservation metadata, provenance | Supports preservation documentation but not holistic repository evaluation |
| This Study | Integrated operational evaluation framework | Expert-validated prioritization | Structured criteria library (Mandatory/Optional) | Governance, system support, workflows, QA/QC mechanisms, access and identifiers | Consolidates functional elements across sources and introduces prioritization for routine and comparative assessment |
This study aims to evaluate repository-level stewardship structures, governance mechanisms, and institutional quality control processes by synthesizing limitations identified in previous frameworks and validating comprehensive evaluation criteria through expert assessment. This framework does not evaluate intrinsic dataset-level scientific quality (e.g., methodological rigor, statistical validity, or completeness). Instead, it assesses repository-level stewardship mechanisms that enable scalable and sustainable quality assurance. By critically analyzing the various requirements presented in existing frameworks and previous studies, this study aims to not only accommodate multiple requirements in the general fields but also set criteria for data repository evaluation. The representative research questions of this study are as follows:
What criteria comprehensively evaluate a data repository?
How relevant and prioritized are these criteria in assessing relevance and priority?
Assessing the relevance and priority of these criteria will help bridge the gap in evaluating data repositories and contribute to a more consistent framework for their management.
2. Research Methodology
In this study, we employed a systematic approach to create and validate criteria for evaluating data repositories. This involved four stages to develop assessment criteria.
Stage 1: Formulation of research question
We initially formulated research questions focused on determining the appropriateness and priorities of various evaluation criteria for the data repository. We investigated existing research data repository evaluation standards and identified limitations.
Stage 2: Case study and analysis
We selected five representative repository-related frameworks and community initiatives, including the RDA, the COAR Community Framework, the Science Europe repository selection criteria, the FAIRsharing registry criteria, and DSpace-based evaluation guidelines. We aimed to identify and document existing evaluation criteria by conducting detailed reviews of their documentation, operational procedures, and published criteria to derive a set of representative evaluation criteria for various repository functionalities and challenges in the next stage.
Stage 3: Empirical investigation with expert validation
This stage emphasized experts conducting an empirical investigation and validation. A survey was conducted with several experts to determine the appropriateness and priority of the derived criteria. The evaluation criteria used Cronbach’s alpha coefficient to ensure consistent internal reliability. Additionally, through priority analysis, based on their average ratings, experts classified elements as essential or optional, providing a balanced framework for repository evaluation.
Stage 4: Framework development through pilot evaluation
We developed a comprehensive framework for data repository evaluation based on the validated criteria in the final stage. The proposed evaluation criteria aim to guide pilot evaluations of data repositories by identifying areas for improvement and additional considerations. We expect this framework to significantly contribute to developing more consistent and effective methods for evaluating and managing data repositories.
3. Case Study and Analysis
To evaluate data storage standards, we applied commonly used evaluation criteria. There are a total of five cases used in this study: functional requirements for RDA, COAR Community Framework for best practices in repositories, repository selection criteria by Science Europe, repository selection and registry criteria by FAIRsharing, and DSpace-based repository evaluation. Since the original text of the evaluation criteria for each case is very large, only the main requirements are summarized here. Full requirements can be found at this link: https://doi.org/10.22747/paper_data.20240711.6.
3.1 Functional requirements of RDA
The Repository Platforms for Research Data Interest Group includes 13 categories and 44 functional requirements. The main requirements of each category are given in Table 2.
Table 2
Main requirements of RDA Data Repository Criteria.
| CATEGORY | REQUIREMENT |
|---|---|
| Metadata | Support for various metadata schemas, including domain-specificity and interoperability, enabling data annotation by owners, authorized individuals, or automatic tools, and evaluating metadata quality. |
| Persistent Identifiers (PID) | Assignment of PID/DOI and integration of PIDs into data management. |
| Authentication | Enable fine-grained authentication and authorization, allow integration or import from external systems, and provide single sign-on with support for various authentication methods. |
| Data Access | Allow data providers to choose the access level (e.g., Open Access), offer state-of-the-art interfaces and clients for the repository’s lifetime, provide authorized users access to data versions, enable embargo date selection, offer sophisticated search capabilities, and allow local downloads of selected information. |
| Policy Support | Enable the automated use of data policies with enforcement points and require all data to be attributed with handling requirements. |
| Publication | Provide data access statistics through external analytics services or internal monitoring, display bibliographic citations for data with export options for citation software, and maintain citations linked to the data. |
| Submission/Ingest/Management | Provide application programming interfaces (APIs) for automated task execution, record audit trails, offer a user-friendly ingest process, ensure integrity and quality control for data and metadata, support micro-services, enable fast data transfer and remote access management, facilitate data and metadata collection with mobile devices, offer definable workflows and both single and batch ingest paths, allow product updates and content deletion by authorized users, provide a vocabulary service, and support staged content. |
| Data Organization | Resource naming through collection virtualization/logical naming |
| Location | Tight integration with (near) data processing. |
| Integration | Support federation, including storage drivers. |
| Preservation and Sustainability | Maintain a permanent history of all data versions, convert files to accessible formats while allowing proprietary and legacy types case-by-case, ensure repository scalability, and support workflows. |
| User Experience/User Interface | Ensure seamless integration of data and research outputs into a coherent discovery and access solution and allow the creation of exceptional collection views or digital exhibitions. |
| Data and Product Quality | Capture ‘degree of confidence’ on each data item. |
3.2 COAR Community Framework for best practices in repositories
The COAR Community Framework provides a global, multidimensional framework that is applicable to various repository types and contexts, divided into essential and desired characteristics. The categories are given in Table 3.
Table 3
Main requirements of COAR Community Framework.
| OBJECTIVE | ESSENTIAL CHARACTERISTICS |
|---|---|
| Discoverability | The repository supports basic and advanced Dublin Core metadata, Open Archives Initiative Protocol for Metadata Harvesting (OAI-PMH) harvesting, tombstone pages for withdrawn resources, PIDs, search functionality, indexing by external services, inclusion in repository registries, and provides metadata in both human-readable and machine-readable formats. |
| Access | The repository offers free access to resources including resource links on landing pages, supports accessibility for persons with disabilities, and provides mechanisms to restrict access to sensitive data for authorized users only. |
| Reuse | The repository includes licensing information in the metadata record, stipulating reuse conditions for the resource. |
| Integrity and Authenticity | The repository employs security practices to prevent unauthorized manipulation, supports metadata revision and resource versioning, and regularly performs integrity checks to detect unauthorized changes or accidental damage. |
| Quality Assurance | The repository conducts lightweight reviews and enhancements of basic metadata upon submission and provides documentation or policies detailing the applied curation processes for resources and metadata. |
| Preservation | The repository has a digital preservation plan detailing management duration, roles, and procedures, records checksums upon submission or modification, collects preservation metadata, ensures that depositor agreements cover necessary preservation actions, allows metadata and resources to be copied or migrated, stores copies in different locations, and maintains a business continuity plan for natural disasters or cyber-attacks. |
| Sustainability and Governance | The repository identifies its managing organization and governance, provides user assistance and dedicated management staff, responds to queries promptly, has a policy for resource management if operations cease, and maintains a long-term management and funding plan. |
| Other | The repository provides public documentation that outlines the scope of the resources accepted in the repository. |
3.3 Repository selection criteria by Science Europe
The repository selection criteria by Science Europe emphasizes selecting reliable repositories organized into four categories (Table 4).
Table 4
Main requirements of Science Europe data repository criteria.
| CATEGORY | REQUIREMENT |
|---|---|
| Provision of PIDs | Allow data discovery and identification, enable searching, citing, and retrieval, and provide support for data versioning. |
| Metadata | Enable finding, referencing, and providing publicly available data, including protected or retracted data, using broadly accepted metadata standards that are machine-retrievable. |
| Data Access and Usage Licenses | Enable access to data under well-specified conditions, ensure data authenticity and integrity, enable data retrieval, provide licensing and permissions information (ideally in machine-readable form), and ensure confidentiality and rights of data subjects and creators. |
| Preservation | Ensure persistence of metadata and data and be transparent about mission, scope, preservation policies, and plans, including governance, financial sustainability, retention period, and continuity plan. |
3.4 Repository selection criteria by FAIRsharing community
The FAIRsharing Initiative presents repository-level and dataset-level criteria included in Table 5.
Table 5
Main requirements of FAIRsharing Initiative.
| CRITERIA | DEFINITION |
|---|---|
| Certification and Community Badges | Does the repository have certification schemes or community badges assessing its fitness, trustworthiness, and adoption? |
| Data Access Conditions | What is the process for requesting and granting access to the repository or dataset, and are the data freely available or subject to a request and approval process? |
| Data and Metadata Standards | What community-defined standards does the repository implement to ensure consistent, machine-readable representation of data and/or metadata, facilitating their discovery and interpretation? |
| Data Curation | Are there minimum curation steps that the repository performs on submitted data, and is there a webpage or document describing the type of curation done? |
| Funding | What type of funding (e.g., grants, donations, memberships) does the repository receive, and which organizations fund it? |
| PIDs for Data | Does the repository assign globally unique and persistent identifiers to the deposited data, and what identifier schema is used? |
| Repository Coverage | What higher-level subject areas, disciplines, and cross-disciplinary domains does the repository cover, including data types, technology, and study? |
| Repository Status | What is the life cycle status of the repository: Is it still being developed, or is it in production and accepting data submissions while possibly undergoing (re)development, enhancement, or maintenance? |
| Resource Sustainability | Does the repository have a webpage or document that describes its sustainability plans? |
| User Support | Does the repository provide a contact point, such as a helpdesk email or contact form, to assist data depositors and users during or after submission? |
| FAIRsharing Record Maintainer | Has the owner or maintainer of the repository claimed the record in FAIRsharing and vetted its descriptions, ensuring verified information and tracking the resource’s evolution, alongside FAIRsharing’s in-house curation? |
3.5 DSpace-based repository platform evaluation guidelines
Tilki et al. (2020) proposed 19 quality criteria for clinical data repositories, which were categorized into the following in Table 6.
Table 6
Main requirements of DSpace-based repository evaluation.
| REQUIREMENT | DESCRIPTION |
|---|---|
| Guidelines for Data Upload and Storage | Supporting various file types and metadata schemas and providing upload mechanisms with instructions. |
| De-identification Practices Before Upload | Providing links to de-identification tools and requirements and implementing de-identification tools. |
| Control of Data Quality | Supporting quality control in its workflow. |
| Formal Contract for Upload and Storage | Incorporating a data transfer agreement in the system workflow. |
| Application of a Metadata Schema | Using a consistent metadata schema, allowing customization, providing tools for metadata completion, and making metadata publicly available. |
| Application of an Identifier | Applying a primary PID system and using other PIDs as appropriate. |
| Flexibility of Access | Allowing open access with optional embargo, web-based self-attestation, managed access through group membership or case-by-case basis, and supporting granular access to different dataset parts. |
| Long-Term Preservation | Supporting long-term data and metadata preservation using sustainable software systems. |
3.6 Synthesis of repository evaluation criteria
We analyzed five existing cases to design comprehensive evaluation criteria for evaluating data repositories. The RDA presents the most comprehensive criteria, encompassing 13 categories, and addresses almost all aspects except certification, community badges, funding, repository coverage, and record maintainers. The COAR prioritizes discoverability, access, reuse, and governance but lacks specific criteria for metadata, authentication, policy support, publication, location, integration, and de-identification. Science Europe emphasizes metadata, PIDs, data access, and preservation but does not cover several other critical areas, such as policy support, publication, and quality assurance. The FAIRsharing community includes a wide range of criteria, featuring unique elements like certification, community badges, and funding, yet it needs detailed criteria for submission, integration, and location. The DSpace-based repository platform evaluation guidelines have specific criteria for de-identification, flexibility of access, and formal contracts, but they need to include broader criteria found in other frameworks.
To address these discrepancies, we proposed comprehensive evaluation criteria integrating the categories and requirements emphasized by these five cases, mainly based on the standard and widely utilized criteria of the RDA (Table 7). A total of 13 requirements and 150 elements were identified among the five cases. Specifically, RDA originally had 44 elements, but due to overlapping elements, we expressed it as 46. Similarly, COAR consisted of 52 elements, but 56 were used in this study for the same reason as RDA.
Table 7
Distribution of evaluation criteria elements across five cases.
| CATEGORY | REQUIREMENT | NUMBER OF ELEMENTS | ||||||
|---|---|---|---|---|---|---|---|---|
| RDA | COAR | SCIENCE EUROPE | FAIRSHARING | DSPACE | TOTAL | |||
| A Data Access and Retrieval | A-1 | Metadata | 3 | 3 | 5 | 1 | 4 | 16 |
| A-2 | Access and Retrieval | 6 | 20 | 3 | 1 | 5 | 35 | |
| B Permanent Identifier | B-1 | Permanent Identifier | 2 | 2 | 3 | 1 | 1 | 9 |
| C Sustainability and Governance | C-1 | Rights, License and Copyright | 1 | 2 | 2 | 5 | ||
| C-2 | Preservation and Sustainability | 4 | 13 | 2 | 1 | 1 | 21 | |
| C-3 | Policy Support and Repository Coverage | 3 | 3 | 4 | 1 | 11 | ||
| C-4 | Repository Certification | 1 | 1 | |||||
| C-5 | Funding | 1 | 1 | 2 | ||||
| D Data Curation and Citation | D-1 | Curation of Data | 15 | 1 | 16 | |||
| D-2 | Data Citation | 3 | 1 | 4 | ||||
| E User Interface and System Support | E-1 | User Interface | 4 | 2 | 6 | |||
| E-2 | System Support | 3 | 4 | 2 | 6 | 15 | ||
| F Quality Control | F-1 | Data Authenticity and Integrity | 2 | 5 | 1 | 1 | 9 | |
The distribution of requirements covered by each case is as follows: RDA and COAR covered 11 requirements each, FAIRsharing covered eight, and both Science Europe and DSpace covered seven requirements each. Notably, the category A-2 Access and Retrieval had the highest number of requirements, with the five cases covering 35 requirements. In contrast, the category C-5 Funding had the fewest requirements, with only two.
4. Empirical Investigation with Expert Validation
The purpose of the survey was to establish evaluation criteria for data repositories and determine the appropriateness and priorities of these criteria. The questionnaire was distributed by e-mail in May 2023 to 40 experts with demonstrated experience in research data repositories, research data management, metadata, and data curation. The participants were identified through existing professional collaboration networks established through previous research activities. They represented universities, government-funded research institutes, private companies, and non-profit organizations involved in repository development, operation, and research data management. Participation was voluntary and all responses were collected anonymously. No personally identifiable information or sensitive personal information was collected. A total of 32 experts responded, resulting in an 80% response rate. The respondents were affiliated with various institutions: 17 from research institutes, 10 from universities, four from private companies, and one from a non-profit organization. Regarding research experience, 16 had more than 20 years, 10 had 10–20 years, four had less than 5 years, and two had 5–10 years. The respondents’ fields included 17 in science and engineering, 11 in humanities and social sciences, three in science, and one who did not respond.
The questionnaire consisted of six categories rated on a 7-point Likert scale: Data Access and Retrieval, Permanent Identifier, Sustainability and Governance, Data Curation and Citation, User Interface and System Support, and Quality Control.
The validity of the response structure was verified by obtaining an index of reliability. Cronbach’s alpha was used to measure the consistency and homogeneity of the items, with values between 0 and 1, where being closer to 1 indicates higher reliability. If k is the number of items in the target item, is the variance of each item, and is the variance of the item to which the item belongs, the Cronbach’s alpha coefficient can be calculated as . The Cronbach’s alpha coefficients for each category are as shown in Table 8.
Table 8
Cronbach’s alpha coefficients for categories.
| CATEGORY | ALPHA COEFFICIENT |
|---|---|
| A Data Access and Retrieval | 0.82 |
| B Permanent Identifier | 0.76 |
| C Sustainability and Governance | 0.94 |
| D Data Curation and Citation | 0.90 |
| E User Interface and System Support | 0.93 |
| F Quality Control | 0.76 |
The results showed that the Cronbach’s alpha coefficient for all categories was over 0.7, indicating that the responses were well-structured and consistent, thus reliable. The survey results contained the importance of each evaluation factor for the data repository evaluation criteria by researchers. Given the high reliability and validity of the survey responses, it was deemed reasonable to establish the data repository evaluation criteria based on these results.
To address the challenge of reflecting the opinions of all researchers in the field, a resampling approach was employed due to the limited number of respondents. By obtaining the sample mean from the survey responses and applying replacement sampling repeatedly, the distribution of the sample mean can be approximated to a normal distribution according to the Central Limit Theorem. The average of these sample means can be estimated as the average response of the entire population of researchers in the related field.
5. Framework Development through Pilot Evaluation
We verified the response averages and priorities of each element of requirements for the comprehensive evaluation criteria for the data repository through empirical investigation. Given that we consistently derive the response average and priority of each element of requirements across all categories, we focus here on the ‘C Sustainability and Governance’ category, which had the highest Cronbach’s alpha coefficients. Detailed information on the response averages and priorities of each element in other categories can be found in the Supplementary file (https://doi.org/10.22747/paper_data.20240711.6).
For each requirement in the ‘C Sustainability and Governance’ category, the priority of the sub-elements was analyzed (Figure 1). The response average and results for each element were graphically displayed, and those with scores lower than the average were designated as optional, while those with scores higher than the average were designated as mandatory (Table 9).

Figure 1
Response score for elements of ‘C Sustainability and Governance’ category.
Table 9
Response average and priority for elements of ‘C Sustainability and Governance’ category.
| REQUIREMENT | ELEMENT | AVERAGE SCORE | OBLIGATION | |
|---|---|---|---|---|
| C-1 Rights, License and Copyright | 1 | Provide information about licensing and permissions ideally in machine-readable form. | 6.38 | Mandatory |
| 2 | Ensure confidentiality and rights of data subjects and creators. | 6.28 | Mandatory | |
| 3 | Allow open access to material, with an optional embargo period. | 5.97 | Mandatory | |
| 4 | Offer managed access through group membership. | 5.16 | Optional | |
| 5 | Support granular access to different parts of dataset collections. | 4.63 | Optional | |
| 6 | The resources in the repository are available at no cost to the user. | 5.34 | Optional | |
| 7 | Provide different access rights for groups and individuals (roles) on collections and allow the import of such concepts (e.g., from identity management systems). In the case of confidential or proprietary data, authenticate every access and authorize every operation. | 5.53 | Optional | |
| 8 | Provide single sign-on and/or support for different authentication methods. | 5.22 | Optional | |
| C-2 Preservation and Sustainability | 1 | Support long-term preservation of data and metadata. | 6.28 | Mandatory |
| 2 | The repository collects basic preservation metadata including provenance, date of upload, and file format. | 5.63 | Mandatory | |
| 3 | The metadata and the resources in the repository can be copied or migrated to other systems. | 5.47 | Optional | |
| 4 | At least one copy of the repository contents is stored in a different location than the original repository. | 5.66 | Mandatory | |
| 5 | The agreement between depositor and repository provides for all actions necessary to meet preservation responsibilities—for example, rights to copy, transform, and store the items. | 5.78 | Mandatory | |
| 6 | Ensure persistence of metadata and data. | 6.41 | Mandatory | |
| 7 | Files need to be converted to the most accessible formats. | 6.28 | Mandatory | |
| C-3 Policy Support and Repository Coverage | 1 | Data policies are used to define what happens when to which dataset. For example, for processing and quality control, regularly enforced policies are helpful. | 5.94 | Mandatory |
| 2 | Policy enforcement points: Control all operations with administrator-defined rules. | 5.68 | Mandatory | |
| 3 | The higher-level subject areas/disciplines that the repository covers, as well as cross-disciplinary domains, such as the types of data, technology, and study. | 4.94 | Optional | |
| 4 | The repository provides public documentation that outlines the scope of the resources accepted in the repository. | 5.16 | Optional | |
| 5 | The life cycle status of the repository: Is it still being developed or is it in production and accepting data submissions? The latter does not exclude that some (re)development, enhancement, or maintenance may be ongoing, as happens in any repository. | 5.19 | Optional | |
| 6 | Be transparent about mission, scope, preservation policies, and plans (including governance, financial sustainability, retention period, and continuity plan). | 5.97 | Mandatory | |
| 7 | The repository has a digital preservation plan that states the duration of time that the resources will be managed for, identifies roles, and documents procedures for the preservation of different resource formats. | 5.72 | Mandatory | |
| 8 | The repository has a business continuity plan that details the response and procedures in case of natural disasters or cyber-attacks. | 6.41 | Mandatory | |
| 9 | Plan that gives information about sustainability plans for the repository: Does the repository have a webpage or document that describes these? | 5.34 | Optional | |
| 10 | The repository provides documentation or has a policy outlining what curation processes are applied to the resources and the metadata. | 5.35 | Optional | |
| 11 | The repository is included in one or more disciplinary or general registry of repositories. | 5.48 | Optional | |
| C-4 Repository Certification | 1 | Certification schemes and/or community badges that assess certain aspects of the repository (e.g., its fitness, trustworthiness, adoption): Does the repository have any? | 5.31 | Optional |
| C-5 Funding | 1 | The type of funding (e.g., grants, donations, memberships) and the organization(s) that fund the repository. | 5.00 | Optional |
| 2 | The repository (or organization that manages the repository) has a long-term plan for managing and funding the repository. | 5.03 | Optional | |
For the C-1 Rights, License, and Copyright requirements, the averages of elements C-1-1, C-1-2, and C-1-3 were 6.38, 6.28, and 5.97, respectively. These scores were higher than the average for the entire C category; thus, we designated these elements as mandatory items. Notably, these elements had an average standard deviation of 1.03, indicating relatively small variability among respondents. Elements C-1-4 and C-1-5 scored below average, so we designated these elements as optional items. For the C-2 requirements (Preservation and Sustainability), most elements scored above average, indicating strong sustainability. For the C-3 requirements (Policy Support and Repository Coverage), there was a mix of high and low scores, suggesting variability in policy support and repository coverage. Elements C-3-3 and C-3-4 scored below average and were designated as optional, while C-3-1, due to its high score, was designated as a mandatory item. The C-4 requirement (Repository Certification) had a score for element C-4-1 that was below average, with a standard deviation of 1.09, indicating low variability and consensus among most respondents. The two elements of the C-5 requirement (Funding), C-5-1 and C-5-2, exhibited high variability and scored below average; thus, both were designated optional items.
6. Proposal of Comprehensive Evaluation Criteria for Data Repositories
The proposed criteria emphasize repository-level stewardship capacity, including governance transparency, preservation planning, integrity safeguards, provenance tracking, workflow control, and sustainable infrastructure. The framework focuses on structural enablers of quality assurance rather than direct scientific validation of dataset content. The comprehensive evaluation criteria comprise six categories, 13 requirements, and 110 elements, including mandatory and optional obligations (Table 10). The ‘E User Interface and System Support’ category had the most elements, totaling 38, while the ‘B Permanent Identifier’ category had the fewest, with only six. This number does not indicate the importance of the categories but rather the number of independent requirements within each category.
Table 10
Comprehensive evaluation criteria for data repositories.
| CATEGORY | REQUIREMENTS | ELEMENTS | SOURCE | OBLIGATION | |
|---|---|---|---|---|---|
| A Data Access and Retrieval | A-1 Metadata | 1 | Support various types of metadata. | RDA | Mandatory |
| 2 | Ensure that metadata are machine-retrievable. | Science Europe | |||
| 3 | Provide information that is publicly available and maintained, even for non-published, protected, retracted, or deleted data. | Science Europe | Optional | ||
| A-2 Access and Retrieval | 1 | Include a link to each resource on its respective landing page in the repository. | COAR | Mandatory | |
| 2 | Enable linking between related contents in the metadata record, such as preprints, published articles, data, and software. | COAR | |||
| 3 | Allow data providers to choose the access level for data (e.g., Open Access). | RDA | |||
| 4 | Enable data or at least metadata retrieval using an open and standardized protocol. | Science Europe | |||
| 5 | Permit local download of selected information. | RDA | |||
| 6 | Support the use of controlled vocabularies in metadata records. | COAR | Optional | ||
| 7 | Facilitate indirect access to restricted resources (e.g., by contacting the author). | COAR | |||
| 8 | Allow data-depositing users to select an embargo date. | RDA | |||
| 9 | Provide a tombstone page for withdrawn resources, ensuring the metadata record remains publicly available. | COAR | |||
| 10 | Offer support to users during or after submission with a contact point (e.g., helpdesk email or contact form) to assist data depositors and users. | FAIR sharing | |||
| B Permanent Identifier | B-1 Permanent Identifier | 1 | Assign a PID/DOI to data and collections during data ingestion or ‘project publication’ time or earlier (e.g., when a paper is submitted but the data are not final yet). Ensure that the PID resolves the research data’s ‘landing page,’ displaying the required descriptive metadata during the embargo period, and establish a clear transition from PID collections to a DOI. | RDA | Mandatory |
| 2 | Integrate all data management activities with PID management. Ensure that PID metadata is always synchronized with the data/metadata holdings. | RDA | |||
| 3 | Ensure that PIDs are included in the corresponding metadata. | Science Europe | |||
| 4 | Consistently assign PIDs (e.g., DOI, URN, ARK) to the data, allowing the corresponding data and metadata to be found, referred to, and retrieved, even if the storage location changes. | Science Europe | |||
| 5 | Support PIDs for authors, funders, institutions, funding programs, and other relevant entities. | COAR | Optional | ||
| 6 | Ensure that metadata information can declare links to other relevant or associated information by providing the PID and a description of the scientific relation, including details of the associated researcher, with permanent research IDs (e.g., ORCID, ISNI, DAI). | Science Europe | |||
| C Sustainability and Governance | C-1 Rights, License and Copyright | 1 | Provide information about licensing and permissions, ideally in machine-readable form. | Science Europe | Mandatory |
| 2 | Ensure confidentiality and rights of data subjects and creators. | Science Europe | |||
| 3 | Allow open access to material, with an optional embargo period. | DSpace | |||
| 4 | Offer managed access through group membership. | DSpace | Optional | ||
| 5 | Support granular access to different parts of dataset collections. | DSpace | |||
| 6 | Ensure that resources in the repository are available at no cost to the user. | COAR | |||
| 7 | Provide different access rights for groups and individuals (roles) on collections, allowing the import of such concepts (e.g., from identity management systems). Authenticate every access and authorize every operation for confidential or proprietary data. | RDA | |||
| 8 | Provide single sign-on and/or support for different authentication methods. | RDA | |||
| C-2 Preservation and Sustainability | 1 | Support long-term preservation of data and metadata. | DSpace | Mandatory | |
| 2 | Collect basic preservation metadata, including provenance, date of upload, and file format. | COAR | |||
| 3 | Store at least one copy of the repository contents in a different location than the original repository. | COAR | |||
| 4 | Ensure that the agreement between depositor and repository provides for all actions necessary to meet preservation responsibilities (e.g., rights to copy, transform, and store the items). | COAR | |||
| 5 | Ensure persistence of metadata and data. | Science Europe | |||
| 6 | Convert files to the most accessible formats. | RDA | |||
| 7 | Ensure that metadata and resources in the repository can be copied or migrated to other systems. | COAR | Optional | ||
| C-3 Policy Support and Repository Coverage | 1 | Use data policies to define what happens to which dataset, enforcing policies regularly for processing and quality control. | RDA | Mandatory | |
| 2 | Implement policy enforcement points to control all operations with administrator-defined rules. | RDA | |||
| 3 | Be transparent about mission, scope, preservation policies, and plans (including governance, financial sustainability, retention period, and continuity plan). | Science Europe | |||
| 4 | Have a digital preservation plan that states the duration of time that the resources will be managed for, identifies roles, and documents procedures for the preservation of different resource formats. | COAR | |||
| 5 | Maintain a business continuity plan detailing the response for various scenarios. | COAR | |||
| 6 | Specify the higher-level subject areas/disciplines that the repository covers, including cross-disciplinary domains, types of data, technology, and study. | FAIR sharing | Optional | ||
| 7 | Provide public documentation outlining the scope of the resources accepted in the repository. | COAR | |||
| 8 | Indicate the life cycle status of the repository: whether it is still being developed or is in production and accepting data submissions. | FAIR sharing | |||
| 9 | Provide information about sustainability plans for the repository, such as a webpage or document describing these. | FAIR sharing | |||
| 10 | Document or have a policy outlining the curation processes applied to resources and metadata. | COAR | |||
| 11 | Ensure that the repository is included in one or more disciplinary or general registry of repositories. | COAR | |||
| C-4 Repository Certification | 1 | Obtain certification schemes and/or community badges that assess certain aspects of the repository (e.g., fitness, trustworthiness, adoption). | FAIR sharing | Optional | |
| C-5 Funding | 1 | Specify the type of funding (e.g., grants, donations, memberships) and the organization(s) that fund the repository. | FAIR sharing | Optional | |
| 2 | Ensure that the repository (or managing organization) has a long-term plan for managing and funding the repository. | FAIR sharing | |||
| D Data Curation and Citation | D-1 Curation of Data | 1 | Maintain a permanent history of versions for all data. | RDA | Mandatory |
| 2 | Ensure that the version of the data stored in the repository is clearly specified and documented via a permanent audit trail for provenance tracing. | Science Europe | |||
| 3 | Support the revision of metadata and versioning of resources. | COAR | |||
| 4 | Incorporate a formal contract regarding upload and storage, including a data transfer agreement in the system workflow. | DSpace | |||
| 5 | Handle staged content, including submission states that are raw, processed, curated, and published. | RDA | Optional | ||
| 6 | Provide an easy-to-use ingest process with minimal barriers to participation. | RDA | |||
| 7 | Define a submission/ingest workflow. | RDA | |||
| 8 | Perform review and annotation of data, ensuring that a set of minimum curation steps are applied to the submitted data. | FAIR sharing | |||
| 9 | Provide a webpage or document describing the type of curation done. | FAIR sharing | |||
| 10 | Register workflows as executable objects and track the provenance of each workflow execution. | RDA | |||
| 11 | Define data access mechanisms and terms at the repository and/or dataset level, outlining the process for requesting and granting access. | FAIR sharing | |||
| 12 | Provide different versions of a dataset. | RDA | |||
| D-2 Data Citation | 1 | Display bibliographic citations for data and allow exporting bibliographic data to citation software (e.g., EndNote, Citavi, Zotero). | RDA | Mandatory | |
| 2 | Ensure that citations provide recognition and updates from others utilizing the data for experiments or other purposes. | RDA | |||
| 3 | Make metadata in the repository available for download in a standard bibliographic format at no cost to the user. | COAR | Optional | ||
| E User Interface and System Support | E-1 User Interface | 1 | Support a responsive, mobile-friendly user interface. | COAR | Mandatory |
| 2 | Provide interfaces (APIs) for automated task execution, such as data ingestion or integration with data analysis tools and other external applications. | RDA | |||
| 3 | Ensure that scientific terms are consistent for future reuse via a vocabulary service. | RDA | |||
| 4 | Provide data access statistics using external analytics services or internal monitoring of user activity. | RDA | |||
| 5 | Offer sophisticated search capabilities for metadata and data, including full-text search and schema-specific search for both humans and computers. | RDA | |||
| 6 | Ensure access to documentation and metadata for individuals. | COAR | |||
| 7 | Allow the creation of special collection views or digital exhibitions. | RDA | Optional | ||
| 8 | Enable fast data transfer, ingestion, and export. | RDA | |||
| 9 | Support data and metadata collection with mobile devices. | RDA | |||
| 10 | Allow authorized users to mark content for deletion. | RDA | |||
| 11 | Collect and share usage information using a standard methodology (e.g., number of views, downloads). | COAR | |||
| 12 | Apply a metadata schema to describe contents and provide tools to help data generators complete metadata fields. | DSpace | |||
| 13 | Allow data annotation by the data owner, authorized individuals, or automatic metadata extraction tools. | RDA | |||
| E-2 System Support | 1 | Support metadata harvesting using OAI-PMH. | COAR | Mandatory | |
| 2 | Implement de-identification practices before data upload. | DSpace | |||
| 3 | Ensure that mechanisms are in place to limit access to authorized users only for sensitive research data. | COAR | |||
| 4 | Recommend tools to anonymize sensitive data to enable data sharing. | COAR | |||
| 5 | Provide mechanisms to make very large files available to users outside the normal user interface when file size becomes unwieldy. | COAR | |||
| 6 | Adhere to the most recent version of the W3C Web Content Accessibility Guidelines. | COAR | |||
| 7 | Ensure that the repository is scalable regarding the amount of data. | RDA | |||
| 8 | Use sustainable software systems. | DSpace | |||
| 9 | Record audit trails to track changes to resource metadata and information relationships. | RDA | |||
| 10 | Integrate data and other research outputs into a coherent and consistent discovery and access solution. | RDA | |||
| 11 | Support both individual and bulk uploads in the repository’s submission system. | COAR | |||
| 12 | Apply a primary PID system. | DSpace | |||
| 13 | Require all data to be attributed with handling requirements, including licenses and security parameters. | RDA | |||
| 14 | Support a range of file types and metadata schema for data upload and storage. | DSpace | |||
| 15 | Provide state-of-the-art user interfaces and clients throughout the repository platform’s lifetime, updating features and functionality to meet current researchers’ requirements and expectations. | RDA | Optional | ||
| 16 | Allow open access after web-based self-attestation of the user. | DSpace | |||
| 17 | Offer managed access through application on a case-by-case basis. | DSpace | |||
| 18 | Encapsulate operations in micro-services that can be chained into a workflow. | RDA | |||
| 19 | Provide both single and batch ingest paths to efficiently submit a range of data types and scales. | RDA | |||
| 20 | Allow product developers to update product information within the repository. | RDA | |||
| 21 | Manage data collections and their properties independently of the storage system and resource naming through collection virtualization/logical naming. | RDA | |||
| 22 | Integrate with near-data processing facilities like High Performance Computing to handle large data volumes efficiently, maintaining provenance information. | RDA | |||
| 23 | Build the repository on well-supported, open-source software. | COAR | |||
| 24 | Map access protocol to storage protocol using storage drivers. | RDA | |||
| 25 | Allow authorized individuals to curate materials from distributed locations through remote access management. | RDA | |||
| F Quality Control | F-1 Data Authenticity and Integrity | 1 | Enforce metadata quality evaluation using metrics. | RDA | Mandatory |
| 2 | Perform regular integrity checks of resources to detect unauthorized changes or accidental damage. | COAR | |||
| 3 | Ensure data authenticity and integrity by including detailed provenance information in the metadata. | Science Europe | |||
| 4 | Support quality control in its workflow. | DSpace | |||
| 5 | Undertake lightweight review and enhancement of basic metadata upon submission of resources. | COAR | Optional | ||
| 6 | Record the checksum when a resource is submitted or modified. | COAR | |||
| 7 | Apply security practices to prevent unauthorized manipulation of resources. | COAR | |||
| 8 | Capture a structured stewardship confidence indicator derived from curation status, validation stage, provenance completeness, and workflow review history, rather than a direct scientific quality score. | RDA | |||
| 9 | Maintain a contact person or organization as the FAIRsharing record maintainer for the repository’s description in FAIRsharing, ensuring that the record is claimed and vetted. | FAIR sharing | |||
The evaluation criteria consist of 110 elements—55 mandatory and 55 optional. Six of the seven elements of the ‘C-2 Conservation and Sustainability’ requirement are mandatory, making up 86% of the requirements. This means that most factors must be taken into consideration in this category. On the other hand, in the requirements of ‘C-4 Repository Certification’ and ‘C-5 Funding,’ all elements are optional. The focus on mandatory elements helps prioritize essential criteria that data repositories must meet to ensure core functionality and compliance. In contrast, optional elements identify additional features that can enhance repository performance and user satisfaction. Mandatory criteria are non-negotiable for operational standards, whereas optional criteria provide valuable enhancements without being essential for basic functionality.
Among the five cases, the elements created based on RDA were the most numerous, with 42 elements. These elements were adopted by COAR, Dspace, Science Europe, and FAIRsharing in the following order: 32, 13, 13, and 11, respectively. This indicates that RDA includes the broadest categories and the most detailed criteria, while FAIRsharing had fewer elements adopted due to less detailed criteria.
7. Discussion
The findings of this study should be interpreted in relation to the established certification and trust-oriented frameworks for digital repositories. Audit-based standards such as ISO 16363 and CoreTrustSeal provide robust and internationally recognized mechanisms for assessing compliance with predefined requirements and organizational controls. These frameworks play a critical role in ensuring repository trustworthiness and long-term preservation governance.
The framework proposed in this study does not aim to replace or replicate such certification mechanisms. Rather, it integrates functional and governance-related requirements extracted from multiple repository-related sources and introduces an expert-validated prioritization mechanism. In this regard, the framework may serve as an operational assessment instrument supporting routine self-evaluation, comparative benchmarking, and preparatory diagnostics prior to formal certification processes.
The distinction between stewardship quality and intrinsic dataset validity is also conceptually significant. The proposed criteria do not seek to evaluate scientific correctness at the content level; instead, they assess infrastructural and governance conditions that enable transparency, traceability, integrity, and reliability within repository operations. Repository-level mechanisms therefore function as structural enablers of trust rather than substitutes for domain-specific validation procedures.
Finally, the introduction of mandatory and optional classifications differentiates this framework from binary certification models. These classifications reflect empirically derived prioritization patterns and may facilitate incremental improvement strategies across heterogeneous institutional and disciplinary contexts.
The framework may also contribute to ongoing international initiatives aimed at standardizing repository service characteristics, including the RDA Community-based catalogue of requirements for trustworthy Technical Repository Service Providers Working Group (TRSPs WG) and related EU-funded projects such as FIDELIS and EDEN, by providing a structured operational synthesis and prioritization perspective.
8. Conclusion
This study aimed to develop a comprehensive repository-level evaluation framework for assessing data stewardship capacity in research data repositories. Effective data management supports data sharing and reuse, both of which depend on adequate dataset-level quality. Because directly assessing the intrinsic quality of all individual datasets and all quality dimensions (e.g., accuracy, completeness, and consistency) is often impractical at scale, evaluating repository-level stewardship capacity provides a structured and scalable approach to strengthening institutional conditions for trustworthy data management. Such evaluation does not replace intrinsic dataset-level scientific quality assessment but enhances confidence in the governance, preservation, provenance, and integrity mechanisms that underpin sustainable data reuse.
This research analyzed five existing repository evaluation frameworks, formulating 13 requirements across six categories encompassing 110 detailed elements. These experts validated these criteria through a survey, confirming their internal consistency, with all items scoring a Cronbach’s alpha coefficient above 0.7.
The survey results highlighted the importance and prioritization of these criteria, categorizing 55 elements as essential (mandatory) and 55 as additional (optional). The six categories identified include Data ‘Access and Retrieval,’ ‘Permanent Identifier,’ ‘Sustainability and Governance,’ ‘Data Curation and Citation,’ ‘User Interface and System Support,’ and ‘Quality Control.’ Each category was designed to evaluate distinct dimensions of repository-level stewardship rather than intrinsic dataset-level scientific validity.
Despite the limitation of relying on five case studies for criteria derivation and the relatively small sample size of experts (survey participants), this study lays a foundational groundwork for data repository evaluation. The proposed criteria can serve as a preliminary framework, fostering further research and refinement through pilot evaluations. Future studies can build upon these findings to enhance and expand the criteria, which can help ensure comprehensive and effective data repository evaluations. This work aims to contribute significantly to the development of more consistent and effective methods for evaluating and managing data repositories, ultimately supporting better data management practices and the advancement of open science.
Ethics and Consent
This study was based on an anonymous online questionnaire survey conducted in May 2023 to obtain expert opinions for validating the proposed repository evaluation criteria. The participants were recruited through existing professional collaboration networks and consisted of experts with experience in research data repositories, research data management, metadata, and data curation. The survey was conducted solely for research purposes. Participation was voluntary, and respondents were informed of the purpose of the study before completing the questionnaire. Completion of the questionnaire was considered to indicate informed consent to participate in the study. No personally identifiable information or sensitive personal information was collected, and all responses were analyzed in anonymized and aggregated form. According to the institutional policy in effect at the time the study was conducted, a formal Institutional Review Board (IRB) approval was not required for this study.
Data Accessibility Statement
The survey data supporting the findings of this study are available from the corresponding author upon reasonable request.
Author Contributions
Sooyeon Han: Conceptualization, formal analysis, writing – original draft, writing – review and editing.
Jong-Gyu Han: Funding acquisition, writing – review and editing.
Sun-tae Kim: Investigation.
Ju-seop Kim: Investigation, methodology, writing – original draft, writing – review and editing.
All authors have read and agreed to the published version of the manuscript.
