Open Science and Technology
The Open Data and Open Science movement is driven by global political interests in getting more transparency in research and the reuse of data for innovation beyond academia. Together with an advancement in data infrastructure, these efforts are expected to provide significant societal benefits and innovation by accelerating data availability and reuse, whilst making research transparent and holding researchers accountable for their work (European Commission, 2025). At the same time, the openness challenges Indigenous needs and interests in data. This is partly because Indigenous data embodies a distinct ontological and epistemological understanding of a phenomenon or worldview. These perspectives are not necessarily captured in the data itself, but are understood by those familiar with the specific Indigenous community. Furthermore, data collected for one purpose may not be suitable for reuse in a different context, as its meaning and validity are inherently context-dependent. The Open Data and Open Science movement, where research data is expected to be openly shared (Bezjak et al., 2018), together with artificial intelligence (AI), requires a paradigm shift in how we control and steward data, and in particular, Sámi and Indigenous data. Considering the ongoing truth and reconciliation processes in the Nordic countries, GIDA-Sápmi calls for Indigenous data sovereignty for research data as a step towards reconciliation.
Indigenous interests intersecting
The GIDA-Sápmi network is an extension of the Global Indigenous Data Alliance (GIDA), a global network of Indigenous and non-Indigenous researchers, data practitioners and activists promoting Indigenous rights to data sovereignty—that is, have data for nation building and the right to govern it (Carroll, Rodriques-Lonebear, and Martinez, 2019). The GIDA network established the CARE principles (Collective Benefit, Authority to Control, Responsibility, and Ethics) (Carroll et al., 2020), which complement the FAIR principles (Findable, Accessible, Interoperable, and Reusable) (Wilkinson et al., 2016). From 2021, GIDA-Sápmi (2022) has promoted the CARE principles in Sámi and non-Sámi scientific communities to secure Sámi interests and control in the collection, use, reuse and storage of research data concerning the Sámi people and Sámi interests, across data ecosystems in Norway, Sweden and Finland. This commentary offers timely perspectives that were raised from the Sámi scientific community on data governance following panel discussions with the titles Sámi Research Governance and Sovereignty and How to Implement Sámi Data Governance and Sovereignty.
The First Sámi Science Week
The inaugural Sámi Science Week, hosted (Sámiráđđi, n.d.) by Sámi Allaskuvla (Sámi University of Applied Sciences) and the Sámiráđđi (Saami Council) in Guovdageaidnu in June 2025, was a significant event where cutting-edge themes across Sámi academia were discussed. The conference committee included two panels dedicated to Sámi data governance, which this commentary—authored by members of the GIDA-Sápmi network—aims to reflect upon. The timing was particularly apt, building on the momentum from the First International Conference on Sámi Research Data Governance held in Romsa (Tromsø) in January 2023 by GIDA-Sápmi and the Centre for Sami Health Research. The conference contributions underscored the importance of implementing the CARE principles in managing Indigenous research data across the Nordic countries, and a relevant theme issue was published in Acta Borealia in late 2024 (Siri and Axelsson, 2024). The Sámi Science Week’s first panel, Sámi Research Governance and Sovereignty, focused on the use of Sámi data for Sámi AI, data restrictions and governance. The panel featured experts representing informatics and Sámi AI, Sámi linguistics, Sámi health data and the GIDA-Sápmi network (2022). At the second panel, How to Implement Sámi Data Governance and Sovereignty, the authors of this commentary gave a short presentation from a variety of research fields. These included Sámi Large Language Models (LLM), Sámi health data, university repository, Sámi perspectives on health data governance and governing bodies, and governance of the data collected by the Norwegian Truth and Reconciliation Commission (TRC).
First panel: Sámi research governance and sovereignty
This panel started with an introduction of the challenges within the Sámi academia, followed by an overview of the CARE and FAIR principles given by the main author of this commentary. The FAIR principles (Wilkinson et al., 2016) are widely adopted principles for research data management promoting Open Science policy. The CARE principles focus on Indigenous data governance (Carroll et al., 2020). They respond to and complement the FAIR principles to make data people- and purpose-oriented, whereas FAIR is about the reuse of data and related technical and compatibility issues. The European Research Council funds excellent and innovative research projects and recognises the need for data to be as open as possible and as closed as necessary (European Commission, 2016) for ethical reasons to protect data. Recently, the Swedish Research Council (2024) included the CARE principles together with the FAIR principles in their recommendations for Good Research Practice in connection with Indigenous research.
The CARE Principles were developed for data ecosystems and for the governance of data. These principles also address how data are used in policymaking through the concept of ‘data for governance’, which emphasises data collection and how data are used to support Indigenous self-determination (Carroll, Rodriguez-Lonebear, and Martinez, 2019). The GIDA-Sápmi network promotes the CARE principles as a means of safeguarding Indigenous and Sámi interests and needs. The principles guide researchers on how to include Indigenous peoples throughout the research process, in the planning of the research, data collection, data storage, data use and dissemination. Moreover, the GIDA-Sápmi network supports the SODA principles (Sámi Ownership and Data Access), which are based on the CARE principles and were adapted and adopted by the Saami Council, an NGO representing Sámi interests in Norway, Sweden and Finland (Fjellheim, 2024). The SODA principles are founded on long-term work (Holmberg, 2021) for promoting pan-Sámi ethical guidelines and reinforcing dignity, authority and the well-being of the Sámi people. These principles are particularly relevant in an era of expanding secondary data use, including AI training and data-driven innovation, which increasingly challenge existing frameworks for Indigenous rights, control and benefit-sharing.
Key comments and remarks to the first panel
There is a Sámi AI lab in progress at Sámi Allaskuvla, providing innovative solutions based on the use of Sámi data and bringing expectations of a digital advancement to the Sámi society. In the discussion, there was a sense of urgency for not taking advantage of the AI technology and a fear of lagging behind, as this is an advanced and fast-growing technology. At the same time, there was concern about the potential misuse of Sámi data when shared openly and reused in AI, potentially together with other registry data, which is common in social and health sciences. Furthermore, the session debated whether Indigenous research was interesting and important beyond the direct benefits to the people studied—that is, Indigenous peoples—or if the Sámi are too few and not interesting enough for the recent technological advancements.
Moreover, the debate touched upon several themes. One was concerning the potential of Sámi language chatbots to support children’s use of Sámi languages, and the risk that the lack of such tools would reduce children’s use of Sámi languages. Another focus was on healthcare applications, such as dementia diagnostic tools, which could become more culturally sensitive and accurate if created by Sámi AI. Other issues concerned the separation of true data from AI-generated data on the Sámi people if data provenance is not traced, the fear of openly sharing Indigenous knowledge, and the potential misuse of it. Despite the variety in the themes debated, there was an understanding that different types of data—ranging from sensitive health data to data on language to Indigenous knowledge and environmental data—need different levels of openness and measures of data governance.
What makes the Sámi AI ‘Sámi’ was unfortunately not discussed. Is it by feeding AI with Sámi data, or by training it to evaluate data and answer prompts with knowledge of Sámi values and historical background, whilst being respectful of Sámi culture? There was limited time to discuss these questions as well as implications of storing Sámi data on cloud platforms outside of Sámi control. Data platforms are often located in different continents and subject to different jurisdictions, further complicating this issue as Sámi data governance guidelines are missing. Many questions remain unanswered—for example, how can Sámi data be leveraged for innovative AI technology that genuinely safeguards our culture, knowledge, lands and people?
Second panel: How to implement Sámi data governance and sovereignty
An interesting notion from this panel was the different perspectives on data governance that was held by Sámi researchers and managers of repositories. Researchers are concerned with the ethics and interests of the individual and the collective—the owners of the data—whereas repositories focus on how to make data available and accessible for reuse in an ethical manner. These different concerns and focuses are addressed by the CARE and FAIR principles, which do not oppose but rather complement each other by decolonising infrastructures holding Sámi data. Local Contexts Notices and Labels were mentioned as possible solutions for digital repositories, as notices enable institutions and researchers to disclose Indigenous interests in data, and labels enable Indigenous communities to reinforce their collective rights and interests to data (Local Contexts, n.d.). Local Contexts balances the FAIR with the CARE principles and addresses provenance, collective ownership and authority of data, and Indigenous knowledge systems, including immaterial rights associated with data. The Labels and Notices are machine-readable tags attached to metadata to communicate rights, responsibilities and governance.
UiT The Arctic University of Norway has developed a grammatical language model that differs from current LLMs, where the language model produces text based upon existing text collections. Unfortunately, there is not much North Sámi text available, representing a narrow genre (bureaucratic text and news dominate; real dialogue is poorly represented) and often containing typographical errors. Instead, at UiT, they have built a model by writing grammar rules in a machine-readable way, resulting in a language model that can analyse any word and sentence, and generate all the wordforms of the language (Moshagen et al., 2024). As such, applying the CARE principles to LLMs may be a question of relevance—to what extent is the model able to help language users, and to what extent does it constitute a threat to the language community? Grammatical models are good at what they do, and their behaviour can be governed by the language community. What they cannot do is generate text. Language models based on neural nets (often called AI models) have shown to give good results for languages with limited resources when it comes to speech technology. For text technology, the picture is more complicated: the resulting text may be fluent, and the word order good, but central (potentially rare) content words are often mistranslated or (especially for the other Sámi languages) even nonwords. In the worst case, generative AI may be counterproductive for small languages and even cause harm to the revitalisation of Sámi languages. Used with care and especially for speech technology, neural models may still become an asset for Sámi and other under-resourced languages.
Research ethics for data collected by the Norwegian TRC and their reuse in research (Broderstad and Josefsen, 2024) have lacked clear ethical guidelines or protocols for the reuse of the data in research (Broderstad and Josefsen, 2023). Some of the core questions raised were: how to protect individuals and small Sámi communities from being identified when reused, how individuals are informed about the reuse of their data for different research projects, and why the TRC’s collection and storage of interview materials were not subject to the same research-ethics standards that researchers must adhere to. Presently, the National Archives of Norway are storing and managing the data, and the reuse of data depends on the Archives’ discretion and judgement, as clear guidelines are still missing. The absence has been evaluated by the National Research Ethics Committee in Norway (NREC, 2024). NREC believes that both the consent form prepared for the TRC’s interviews and the guidelines for access to the TRC’s archives appear unclear and somewhat incomplete, but this does not constitute a research ethics obstacle to access in itself. However, they recommend the use of an ethics checklist to create awareness of the sensitive characteristics of the data. From Broderstad and Josefsen’s (2023) perspective, it is essential to focus more on a responsible use of data to address the interests and needs of the people who have provided the data, and the application of the data needs to be in accordance with sound ethical standards. Collectively, ensuring that TRC data are used responsibly could provide a first step towards Sámi data governance.
Eleven focus group interviews were conducted with Sámi individuals in Sweden (Axelsson and Storm-Mienna, 2021). One of the questions posed was: Who should own or govern Sámi health data? The answers offered solutions to different elements of data governance. The focus groups considered good data management of Sámi research data as important for building trust and acceptance for research. Regarding ownership, the focus groups suggested healthcare authorities, universities, and the Sámi Parliament as owners, wherein the majority considered the last one as the most appropriate owner. However, there were doubts about whether the Sámi Parliament had the capacity to govern Sámi data. If data were governed outside the Swedish Sámi Parliament, decisions regarding the use and reuse of Sámi data should only be made by people who are knowledgeable about Sámi culture, living conditions and history. This was regarded as necessary, as misuse of data and state power still was present in the collective consciousness of the Sámi people. Interestingly, the Sámi focus groups call for Sámi governance of Sámi data and acknowledged capacity—for instance, funding for infrastructures and governance mechanisms with review boards—as a significant barrier to achieving Sámi data governance.
The Population-Based Study on Health and Living Conditions in Regions with Sámi and Norwegian populations, the SAMINOR Study, has a review board that oversees the use of SAMINOR data. This ensures that research projects comply with individual consent and the overall aim of the study, and that projects are conducted in collaboration with Sámi knowledge holders and approved by the Regional Committees for Medical and Health Research Ethics (REC). Additionally, the Norwegian Sámi Parliament has appointed a review board that ensures that health research relevant to the Sámi people, their languages, culture or societies is conducted in partnership with the Sámi people, according to Sámi values, and, if approved, this board gives a collective consent on behalf of the Sámi people. Together, the SAMINOR project board and the Sámi Parliaments’ review board align well with the aims of the CARE principles and exemplify how the CARE principles may be operationalised for sensitive data (Siri, Melhus and Broderstad, 2024). These boards oversee that collective benefits are fulfilled, and that research is ethical and responsible. Whilst the boards oversee that the Sámi people have the authority to control, this control is not absolute, as collective Sámi ownership of data is not acknowledged. Instead, under the General Data Protection Regulation and the Norwegian Personal Data Act (2018), the data hosting university (UiT The Arctic University) serves as the data controller or owner of the data from the SAMINOR Study.
Key comments and remarks from the second panel
In the following discussion, members of the audience emphasised the need for practical examples that demonstrate how the CARE principles can be applied, and how Sámi data governance can strengthen the trust between academia and Sámi society. Collaboration between Sámi academia and the respective Sámi Parliaments in Norway, Sweden and Finland was posed as a potential starting point. This would position the Sámi Parliaments as gatekeepers of the Sámi data and provide an example for how to operationalise Sámi data governance.
The audience valued the CARE principles for advancing the critique of data-holding institutions that do not address collective rights. Also, the CARE principles were regarded as important as they emphasise the need for cultural understanding, Sámi partnership, and a clearly defined ownership of data before it is collected, used and reused. As such, it was suggested that CARE and FAIR principles should be included in the curriculum for those in academic training. Moreover, the audience found the CARE principles valuable as they hold researchers more accountable for doing research that is actually beneficial to Sámi society and Sámi interests, and direct users to Sámi research ethical guidelines.
Furthermore, the panel suggested that the national-specific Sámi Parliament could appoint a committee to oversee Sámi research data governance. For the Norwegian context, it was suggested to extend the mandate of the Sámi Parliaments’ review board to include data governance. Additionally, it was suggested that the national-specific Sámi Parliaments should formulate an explicit set of founding Sámi research data governance principles, with clear protocols. Without founding Sámi principles and protocols, it is challenging to know what to govern by, and how to conduct responsible and ethical research. Moreover, the panel agreed that Sámi data governance could rely on the SODA (Fjellheim, 2024) and CARE principles, with an acknowledgement that the level of data governance may vary by research discipline. Importantly, the panel suggested that Sámi academia should be encouraged to apply the CARE principles in data governance, and through that, test and offer practical solutions on how these principles may be operationalised for Sámi data within the various disciplines. Founding Sámi data governance principles, or the CARE or SODA principles, may ease data governance for researchers collecting the data, and for archives, museums, and institutions hosting Sámi data. It was clearly stated that the Sámi academia cannot solve how data governance should be, but can, through research, contribute to mapping what the Sámi people want and offer solutions for how this can be done. As the national-specific Sámi Parliaments are the highest governing Sámi body for the Sámi people within their respective states, they have the legitimacy to formulate the founding Sámi data-governing principles and protocols. The forthcoming pan-Sámi ethical guidelines by NordForsk (2026) might give some answers regarding founding the Sámi principles for data governance. These may also contribute to easing the sharing of Sámi research data across national borders within Sápmi, which is currently challenging, given the different national policies and the General Data Protection Regulation.
On the other side, there were concerns regarding whether a data governance committee appointed by a Sámi Parliament gives political control over research, potentially influencing how research is conducted and under what conditions. Another concern was that appointing a data governance committee could introduce stricter policies that place additional administrative burdens on researchers, requiring them to navigate across multiple steps to obtain approvals of research protocols. For example, the use of SAMINOR data, which requires approvals from the SAMINOR project board, the Norwegian Sámi Parliaments’ review board, in addition to the REC, was perceived as redundant since it caused delays in research. Collectively, these concerns reflected apprehension that an additional governance review board could ultimately tier researchers and reduce the volume of research undertaken. Regardless of how data governance is solved, the panel unanimously supported a set of founding Sámi data governing principles.
Finally, there were discussions as to whether Sámi data in the future should be stored in one infrastructure or institution to maintain a high level of governance, or if data should remain stored in the various institutional infrastructures but be governed by Sámi founding data governance principles and protocols. There are already Sámi institutions, such as museums and archives across Finland, Sweden and Norway, that host Sámi data. If these institutions were given the authority by Sámi founding principles to govern the data, it would be expected that the use of Sámi data would become more beneficial to the broader Sámi societies.
To Summarise
Even if the concept or discourse of the CARE principles and Indigenous data governance is fluid, sometimes fragmented and subject to ongoing debate, the panels concluded that there is a need for Sámi governance over Sámi data. To address this, we need Sámi governing principles and protocols that can guide the collection, storage, use and reuse of data, whilst recognising that the level of governance may vary according to data and research discipline. It was suggested that the first step towards Sámi data governance would be the establishment of Sámi founding principles that can guide researchers and data-holding institutions. Researchers, in particular, must consult with participants—both individually and collectively—to determine their preferences regarding restrictions on data use, storage and reuse. Implementing Sámi-specific governance principles could foster more responsible data management and rebuild trust within Sámi communities, which remains fragile due to historical misuse of research data for assimilation and racial superiority claims. As the capacity for data storage within Sámi institutions may be challenging, a solution could be a co-ownership model between the respective Sámi Parliaments and the data-hosting institution. Institutions and data archives need to open their research data governance to include Sámi interests and needs, which would contribute to increasing the trust and relevance of research. Presently, the Ethical Guidelines for Research Involving the Sami People in Finland (Heikkilä et al., 2024), together with the SODA principles, may act as the Sámi founding governing principles, as they provide guidelines for ethical research with Sámi societies and guidance on Sámi data governance.
One year after these panels took place, the discussions remain highly relevant. The rapid development of AI has further highlighted the importance of Sámi data governance. However, initiatives such as GIDA’s ongoing work on AI and Indigenous Data Sovereignty and the establishment of the Sámi AI Lab in Guovdageaidnu demonstrate a shift from discussion to practice. Currently, the Sámi AI Lab focuses on community outreach activities where engagement with Sámi communities defines and shapes the use of the technology. We thank the conference committee for bringing this topic forward in Sápmi by devoting space for the panels. We hope that Sámi academia, museums and archives alike, together with the Sámi Parliaments and Sámi society, will contribute to further advancement of Sámi data governance to achieve better data for governance.
Author Contributions
The GIDA-Sápmi network had a panel at the first Sámi Science Week held in Norway, Guovdageaidnu/Kautokeino, in June 2025, and participated in two events, as described in this commentary. SRAS and PA wrote the abstract and main text. All panellists of the second panel described their work and approved this commentary, that is, SEG specified the policies around university library repositories, TR and SNM described their work with Sámi language models, EGB and EJ wrote about governance of the data collected by the Norwegian Truth and Reconciliation Commission (TRC), CSM summarised findings from Swedish Sámi focus groups on data governance of Sámi health data, ARB on governance of the repeated and longitudinal SAMINOR Study, MJH helped structure the abstract with RK, who also moderated the second panel and was a member of the organising committee to Sámi Science Week. All authors contributed to the critical review and revision of the manuscript, and approved the final version for publication.
