Skip to main content
Have a personal or library account? Click to login
Automated Retrieval and Analysis of Finnish Cultural Metadata with finna Cover

Automated Retrieval and Analysis of Finnish Cultural Metadata with finna

Open Access
|Jul 2026

Full Article

1 Introduction

Cultural heritage refers to both material and immaterial attributes that are inherited from past generations, maintained in the present, and preserved for future generations (UNESCO, 2009; Vecco, 2010). The Web recently emerged as a major platform for publishing digital cultural heritage content, and its role in preserving, sharing, and analyzing cultural heritage has become increasingly important. Systematic metadata, i.e., descriptive information about these attributes, and paradata, i.e., information on the processes and tools involved in creating digital cultural heritage objects (Ioannides et al., 2025), are increasingly available in digital form at large scale. Their availability can support discovery and contextualization (Dooley and Bowers, 2018; Riley, 2017; Umerle et al., 2022). For instance, bibliographic catalogues describe the properties of printed books, such as titles, authors, genres, languages, dimensions, and other details. The value of such systematic resources for research has been recognized for decades (Tanselle, 1974). Digitalization has created new means to collect, share, and utilize this information for research purposes (Kruusmaa et al., 2025; Lahti et al., 2019; Padilla et al., 2019; Tasovac et al., 2020; Umerle et al., 2022).

Over the past decade, cultural heritage institutions have increasingly improved access to and reuse of digital collections. Galleries, Libraries, Archives, and Museums (GLAMs) are now routinely publishing structured metadata via application programming interfaces (APIs), enabling programmatic access beyond traditional catalog interfaces, such as search portals and other systems designed for interactive human use (DPLA, 2024; Europeana, 2024; Tønnessen and Birkenes, 2025). The international GLAM Labs community and projects such as the GLAM Workbench have complemented data infrastructures with practical tools (Gustavo Candela et al., 2026). Whereas these platforms have enabled a wide range of research, metadata heterogeneity, data quality, and the sustainability of methods are posing general challenges in the digital cultural heritage research (Mäkelä et al., 2020; Umerle et al., 2022). This is particularly evident in national aggregation platforms, where metadata is contributed by diverse institutions following different cataloging traditions and standards (Tasovac et al., 2020; Umerle et al., 2022). Heterogeneous sources, uneven data quality, and the lack of statistical research tools impose technical burdens on research-oriented computational analysis. Moreover, the tools are often project-specific and face challenges related to long-term maintenance and integration into data analytical environments. The diversity of data sources and formats challenges the integration of the resources into reproducible research workflows. In response, national libraries in Europe have increasingly adopted “Collections as Data” approaches, providing curated metadata, persistent identifiers, and access to digitized collections to support digital scholarship (Gianfranco Candela et al., 2024; Padilla et al., 2019). For example, the National Library of Scotland has developed the Data Foundry data-delivery platform, which publishes digitized collections along with enriched descriptive, technical, and preservation metadata (National Library of Scotland, 2025), and the National Library of Norway has established a national metadata infrastructure for bibliographic and archival metadata (National Library of Norway, 2021). European initiatives such as the European Literary Bibliography (ELB) aim to harmonize and enrich multilingual bibliographic metadata drawn from national and specialist bibliographies (Institute of Czech Literature and Czech Academy of Sciences, 2024).

Despite growing interest in the computational reuse of cultural heritage metadata, there are few actively maintained open-source software packages that support reproducible statistical research through programmatic access to cultural heritage APIs. A small number of R and Python packages have been developed for this purpose for cultural heritage collections. Examples of R packages include rdpla, an interface to the Digital Public Library of America (DPLA) (Chamberlain, 2016), and europeanaR, which provides access to the Europeana REST API (Kourouklis, 2017). In the Python ecosystem, corresponding tools include pyeuropeana, a comprehensive wrapper for the Europeana APIs (Europeana, 2023), the DPyLA client for accessing DPLA metadata (DPLA, 2022), and institution-specific tools such as the Smithsonian Open Access API Python client (Smithsonian Institution, 2023). Beyond direct API clients, archdata provides archaeological datasets for quantitative analysis (Carlson and Roth, 2021). While such tools demonstrate the research value of cultural heritage metadata, several are no longer actively maintained or remain limited in scope. Moreover, many function primarily as thin wrappers around HTTP endpoints, leaving users responsible for managing API constraints, pagination and batching strategies, metadata heterogeneity, and downstream harmonization across collections. This places a substantial technical burden on researchers and has limited adoption outside technically specialized communities (Lahti et al., 2019). Despite the widespread use of shared metadata standards, research-oriented methods workflows in open statistical programming environments such as R remain scarce. Thus, there is a lack of sustainable, research-oriented software within the current landscape that integrates cultural metadata into widely used analytical environments while addressing reproducibility, data quality, and workflow automation. We argue that these problems can be at least partially overcome and sustainability of the software enhanced by using standardized software package development guidelines and established open research software repositories such as the Comprehensive R Archive Network (CRAN).

The lack of reproducible tools to support statistical research on Finnish cultural heritage motivated our work, and the need for hands-on access to data and statistical programming tools for quantitatively oriented researchers has guided the design of the finna package (Tolonen et al., 2019a). Here, we demonstrate how to bridge Finnish cultural metadata with the R programming ecosystem through the open-source R package finna. It interfaces with the Finna platform, aggregating metadata from Finnish libraries, archives, and museums. Two brief case studies are used to demonstrate the approach. Our work aligns with established practices in data science and digital humanities (Mäkelä et al., 2020; Wickham, 2014), and it can support cultural heritage research in Finland and demonstrates the wider utility of the approach for building reproducible research tools for studying cultural heritage research. In addition, the project provides new data resources for methodological research and education in applied statistics and data science.

2 Data Resources

Finna is a digital platform that provides access to a vast collection of cultural and scientific materials from GLAM institutions across Finland. It was developed by the Finnish Ministry of Education and Culture as a unified search and access service that allows users to discover materials such as books, articles, photographs, maps, audiovisual materials, and various artifacts. The platform integrates data from a wide range of Finnish institutions, regional archives, university libraries, and cultural heritage organizations and is accessible at https://www.finna.fi.

Finna provides a search interface supporting keyword, author, title, and subject queries, with filters for format, language, and contributing institution. It has user interfaces and metadata available in multiple languages, including Finnish, Swedish, and English. The platform has an open API, which allows developers to build applications and researchers to develop statistical research workflows. The API provides access to metadata for altogether more than nine million records available under an open license (CC0), enabling reuse by researchers and the general public (Laine and Virtanen, 2016; National Library of Finland, 2023).

2.1 Coverage

At the time of writing, Finna aggregates metadata from hundreds of contributing institutions, including national, university, and public libraries, museums, and archival organizations. The platform provides access to tens of millions of metadata records describing a wide range of material types, such as books, serials, manuscripts, sheet music, sound recordings, audiovisual materials, images, and collection- and item-level archival descriptions. Archival materials are explicitly represented within Finna; however, archival descriptions vary in granularity and typically do not include full-text content. Coverage across institutions and material types remains uneven, and not all holdings of participating organizations are accessible through Finna due to copyright restrictions, licensing constraints, or institutional policies. While Finna provides broad national coverage, certain institutions and material types are missing; smaller local archives, private collections, and community-based heritage initiatives do not always expose their metadata through the platform. In addition, Finna primarily aggregates descriptive metadata; born-digital research data, administrative records, and institutional operational datasets are generally outside its scope (Table 1).

Table 1

Overview of the Finna service and metadata coverage.

ASPECTCOVERAGE
Participating organizationsMore than 500 Finnish libraries, archives, museums, and other cultural heritage organizations.a
Material typesBooks, journals, manuscripts, sheet music, sound recordings, audiovisual materials, images, archival descriptions, maps, museum objects, and electronic publications.
Access methodsFinna web interface, REST API, and OAI-PMH services.
Searchable metadata fieldsMore than 18 searchable metadata fields are available through the Advanced Search interface, including title, author, subject, description, classification, identifier, publisher, publication year, and media type.b
Metadata richnessOAI-PMH harvesting supports multiple metadata schemas (e.g., MARC21, Dublin Core, LIDO, and EAD), ranging from relatively compact schemas with fewer than 20 core elements (e.g., Dublin Core) to highly detailed schemas containing hundreds of metadata fields and subfields (e.g., MARC21).
LimitationsCoverage varies across institutions and collections; some records and digital objects are restricted by copyright, licensing, or institutional policies.

[i] a Complete list of participating organizations: https://www.finna.fi/Content/organisations.

b Complete list of searchable metadata fields: https://www.finna.fi/Search/Advanced.

Finna contains several core Finnish data collections maintained by the National Library of Finland and partner institutions, offering researchers unified access to bibliographic, archival, and audiovisual resources. Among the prominent collections is Fennica, the Finnish National Bibliography, which documents the nation’s literary heritage since the fifteenth century and includes monographs, maps, audiovisual materials, and electronic publications printed in Finland, as well as publications related to Finland that are printed outside Finland. The Viola collection contains data on the Finnish national discography and sheet music bibliography, providing valuable insights into Finland’s musical heritage. It also contains foreign materials in 13 Finnish music library collections (National Library of Finland, 2025c). Ephemera collection consists of printed ephemeral materials published since the nineteenth century, documenting the activities of organizations and institutions through materials such as guides, exhibition catalogues, timetables, brochures, reports, advertisements, invitations, price lists, and telephone directories. Humaniora includes Finnish and international literature in the humanities and social sciences; in addition to research literature, it contains a substantial number of reference works, bibliographies, and primary source publications covering fields such as history, art, philosophy, linguistics, and theology. Beyond these core collections, Finna also integrates additional specialized resources, including the Slavonic Library (Slavica), the manuscript collection, Finnish article databases such as Arto and Journal.fi, the National Repository Library (Vaari), and printed holdings accessible via Helka. Together, these resources substantially broaden the scope of materials available for cultural and historical analysis (Table 2) (National Library of Finland, 2025a).

Table 2

Major collections accessible through the National Library’s Finna service.

COLLECTIONDESCRIPTIONRECORDS
VaariNational Repository Library holdings, including monographs, serials, dissertations, and Braille materials.10,480,618
ArtoBibliographic database of Finnish journal and monograph articles.2,328,020
FennicaFinnish National Bibliography covering books, serials, maps, audiovisual materials, and electronic publications relating to Finland.1,323,790
HelkaResearch library collections from the University of Helsinki and partner institutions.1,232,098
ViolaFinnish national discography and national bibliography of sheet music.1,162,875
HumanioraHumanities and social sciences research literature and reference collections.766,898
SlavicaSlavonic Library collections supporting research on Russia and Eastern Europe.255,026
Journal.fiFinnish scholarly journals and annual publications.103,362

[i] Collection descriptions and record counts were obtained from the National Library of Finland collection description and collection search services (National Library of Finland, 2026a, b).

2.2 Limitations

Access to some archival materials is limited by legal restrictions and institutional access policies. Sensitive records, personal data, and materials subject to access restrictions are often excluded or described only at a high level. Likewise, many copyrighted works, especially recent audiovisual materials, commercial sound recordings, and digitized full-text books, are not available as openly accessible digital objects, even when their metadata records are present. As a result, Finna should be understood as a discovery and metadata aggregation service rather than a comprehensive digital repository of all Finnish cultural heritage holdings.

Finna aggregates metadata originating from heterogeneous source systems that use different cataloguing traditions and metadata standards. Library records are commonly based on bibliographic formats such as MARC21, while museum and archival descriptions rely on domain-specific schemas and descriptive practices. During aggregation, these heterogeneous records are mapped into a shared, discovery-oriented data model optimized for search and retrieval rather than full semantic harmonization. As a result, core metadata elements such as identifiers, titles, creators, dates, languages, formats, subjects, and institutional provenance are generally available across records, but their completeness, structure, and interpretation may vary substantially between institutions and collections (Dooley and Bowers, 2018; Riley, 2017). The use of controlled vocabularies and classification systems is similarly uneven. Subject headings, genres, and classifications may originate from national authority files, institutional thesauri, or free-text descriptions, depending on the contributing institution and historical cataloguing practices. Finna exposes these elements as provided by the source systems and does not enforce full normalization across collections. Consequently, Finna should be understood as a large-scale aggregated research resource whose metadata reflects the diversity and historical changes of Finnish cultural heritage (Mäkelä et al., 2020; Umerle et al., 2022).

2.3 Data standards and structures

Finna primarily uses YSO (General Finnish Ontology) for subject indexing. This is a Finnish-native ontology maintained by the National Library of Finland based on the Simple Knowledge Organization System W3C standard. For shelf-classification most public libraries use YKL (Finnish Public Libraries Classification System), a Finnish adaptation of Dewey Decimal Classification, whereas academic libraries typically use the Universal Decimal Classification (UDC). The underlying bibliographic format is MARC21. Since Finna aggregates collections from many institutions, multiple classification schemes can appear in search results.

The fields are passed through as provided via the Finna APIs (National Library of Finland, 2023). Finna offers two complementary access mechanisms: a REST-based API designed for flexible querying and exploratory access, and an Open Archives Initiative Protocol for Metadata Harvesting (OAI-PMH) harvesting interface intended for systematic bulk retrieval. The REST API supports keyword search, faceted filtering, and record-level access, while the OAI-PMH interface enables harvesting based on predefined sets and incremental updates but provides more limited filtering capabilities (National Library of Finland, 2023). This is designed for systematic harvesting of large metadata collections using predefined sets and incremental updates. However, it offers limited filtering capabilities and does not support the same level of query expressiveness as the REST API (Lagoze and Van de Sompel, 2002).

3 Methods

The finna package addresses three key challenges: automated data retrieval from Finna-API, metadata refinement and quality control, and integration with R workflows. Let us next summarize the overall capabilities of the finna package.

3.1 Data retrieval and API integration

The finna R package builds on both the REST API and OAI-PMH interfaces to support research-oriented workflows. It does not replicate the Finna APIs; instead, it provides higher-level abstractions that automate batching, manage API constraints, standardize outputs into tidy data structures, and support reproducible retrieval strategies. By integrating both interfaces, the package allows researchers to select the access mechanism best suited to their analytical goals while working within a consistent R-based workflow (Lahti et al., 2019).

The package facilitates the systematic collection of large datasets through automated batching strategies that can overcome API-imposed record limits. Records can be retrieved by, e.g., dividing queries into year-based ranges while explicitly including records with missing or undefined years to avoid systematic data loss (Lahti et al., 2019; National Library of Finland, 2023). Moreover, the package supports the inclusion of so-called hidden records that are excluded from standard searches via the Finna user interface but are accessible through the API. These may correspond to restricted material or specialized institutional collections. The data composition can be explored by applying filters such as ‘collection:“FEN””’ and ‘collection:“VIO””’ to focus on the Fennica and Viola collections, respectively. The flag ‘“finna.include_hidden_parts:1”’ enables the user to access hidden collections and items not visible to the public user interface. These filtering options enable users and researchers to obtain a complete view of available metadata. The retrieved data can be analyzed directly or combined with other sources to further enrich and link semantics in analytical pipelines.

The Finna-API provides structured, programmatic access to extensive cultural metadata collections, enabling integration with computational environments such as R. Despite extensive API documentation, integrating these resources into analytical workflows remains a challenge for researchers who rely on R for data analysis. This has been due to the lack of customized methods for retrieving data in a directly analyzable format for automated, reproducible computational workflows. The finna package bridges this gap by offering an intuitive interface to query the Finna-API and facilitating the direct import of metadata into R for further processing and visualization.

3.2 Metadata refinement and quality control

Cultural heritage metadata aggregated from multiple institutions inevitably exhibits variation in completeness, structure, and quality. In the Finna platform, such inconsistencies arise from differences in cataloging practices, metadata standards, and historical documentation processes across contributing libraries, archives, and museums (Dooley and Bowers, 2018; Riley, 2017). Common issues include missing publication years, incomplete language information, heterogeneous format labels, and uneven use of controlled vocabularies (Lahti et al., 2020; Tolonen et al., 2019a). Practical metadata quality assessment in large-scale, aggregated bibliographies typically emphasizes the systematic detection of incompleteness, structural inconsistencies, and schema-level issues rather than full semantic harmonization (Király, 2023).

The finna package addresses these challenges during data retrieval, and its integration into the R ecosystem enables further customization to specific research questions (Mäkelä et al., 2020). The automated batching strategies help to avoid error-prone manual intervention and including records with missing or undefined years will enhance coverage and reduce selection bias (Lahti et al., 2020; Tolonen et al., 2019a).

The package also supports automated curation and refinement by facilitating the direct integration of metadata into tabular data science workflows (Lahti et al., 2019; Wickham, 2014) by transforming the retrieved records into tidy, tabular data structures suitable for statistical analysis in R. Visualization methods are provided, e.g., illustrating top entries and longitudinal trends for quick data exploration and quality control.

Although the finna package is primarily designed for data retrieval, it also provides methods for data refinement and visual exploration to facilitate curation. In this capacity, it remains more limited since the interpretation of missing values, resolution of semantic inconsistencies, and domain-specific harmonization remain the responsibility of the researcher and are often tied to the specific research questions, requiring collaboration with domain experts such as librarians or curators (Mäkelä et al., 2020; Umerle et al., 2022). The package lowers the technical barrier to data access while making data quality issues explicit and traceable within reproducible workflows. Compatibility with other related packages (e.g., the fennica R package (Lahti et al., 2016)) further supports data quality control, validation, and integration into broader bibliographic research workflows. In addition, the finna package can be combined with complementary tools, such as the finto R package, to support metadata enrichment, including author information through linkage to related authority data sources such as Kanto. This enables standardized author identifiers, preferred name forms, and background information to be associated with Finna records, supporting author disambiguation and semantically consistent downstream analysis.

3.3 Package implementation and design

The finna package addresses the need for automated, large-scale, and reproducible access to heterogeneous Finnish cultural heritage metadata by providing a user-friendly R-based interface to both the Finna-API and the OAI-PMH harvesting interface, supporting comprehensive metadata retrieval and analysis. This package facilitates efficient access to Finna’s rich metadata, enabling researchers to use automated, reproducible computational workflows for metadata-driven analysis. It enables researchers to integrate Finna resources into R-based workflows for large-scale metadata analysis (Figure 1).

Figure 1

Schematic diagram for finna package and Finna-API interaction. The package retrieves metadata records from the Finna-API and supports downstream processing in the R statistical environment. The package enables structured, reproducible access to Finnish cultural heritage metadata.

The Finna-API provides complementary functionalities for accessing and enriching Finnish cultural metadata Figure 1. Finna-API facilitates various data interactions, including searching records, retrieving specific records by ID, and retrieving public lists of records. The package interfaces with the primary Finna-API endpoints (/search, /record, and /list) and supports OAI-PMH harvesting for large-scale metadata retrieval, allowing users to systematically collect large-scale cultural metadata in accordance with standard metadata harvesting protocols widely used in libraries and archives. The finna package is designed to interface with the Finna-API, providing seamless access to metadata from Finnish cultural institutions, such as libraries, archives, and museums. The key API endpoints /search, /record, and /list allow users to perform complex searches, retrieve specific records, and access curated public lists (National Library of Finland, 2023).

The finna package is implemented in R (version ≥ 3.5). It exploits widely used libraries for data processing, metadata handling, and visualization, including dplyr, httr, jsonlite, xml2, ggplot2, and purrr. Development adheres to established R software engineering standards, with modular functions, automated testing using testthat, and comprehensive documentation. The package vignette provides practical examples and detailed guidance on querying the Finna-API, managing search parameters, and integrating metadata into reproducible R workflows. These features ensure that researchers can efficiently access and analyze Finnish cultural metadata in alignment with open science principles (Lahti et al., 2017).

Based on good practices in bibliographic and data science (Wilson et al., 2017), finna R package adopts a modular design with reproducible data retrieval and harmonization workflows. This approach aligns with emerging standards in metadata processing that emphasize reproducibility, automated enrichment, and data quality monitoring, as seen in large-scale harmonization efforts across European collections, including Finnish, Swedish, and other national bibliographies (Lahti et al., 2019; Tolonen et al., 2019b).

3.4 Software availability

The finna package is available on Zenodo at https://doi.org/10.5281/zenodo.15655377. The development version is also openly accessible under a BSD-2-Clause license, where users can explore the source code, contribute to the development, and report issues (Morin et al., 2012). Comprehensive documentation and example use cases for finna can be found in the package vignette, which guides users in querying the Finna-API, utilizing search parameters and integrating the retrieved metadata into analytical workflows. Similarly, reproducible scripts for the use cases presented below can be found at https://github.com/fennicahub/finna. These resources enable users to leverage the package for in-depth research on Finnish cultural heritage.

4 Applications

Having described the finna package architecture and functionality, let us next demonstrate its utility through two case studies examining Finnish bibliographic and music metadata.

4.1 Use cases

The first use case involves the collection and analysis of Fennica metadata, with a focus on a specific and historically significant event. The second use case explores Viola, which was similarly collected and analyzed.

4.2 Case study 1: Finnish national bibliographic metadata

Spanning books from as early as 1488 to contemporary materials, Fennica documents not only publications printed in Finland, but also works about Finland or authored by Finnish writers abroad (Marjanen et al., 2019). With more than 1.2 million monograph records and thousands of other entries ranging from maps and audiovisuals to ephemera, the database reflects Finland’s cultural and intellectual change over centuries (National Library of Finland, 2025b). We use finna package to explore the Fennica data collections and study the history of Finnish literary production, which includes the number of books published, topics covered, and languages used.

We present a descriptive bibliometric analysis of Finnish language publishing between 1809 and 1917 based on Fennica metadata retrieved via the finna package. The analysis examines publication volume, language distribution, publication formats, metadata completeness, and institutional participation over time. We explore patterns in Finnish-language publishing as reflected in the data from 1809 to 1917, when Finland was an autonomous Grand Duchy under Russian rule (Launis et al., 2025). To address this, we implemented shortcut functions in the finna R package to programmatically retrieve all bibliographic metadata from the Finna-API. We use the query filter ‘collection:“FEN”’ and the function by default includes the hidden records and undated entries to ensure comprehensive coverage. In total, we retrieved about 65,870 records. The complete record was collected in about 6 minutes using the batch query and a regular laptop.

The language frequency analysis reflects the steady increase of the share of the Finnish language, as compared to Swedish (Figure 2). The growing share of Finnish-language publications reflects broader societal developments, including the rise of Finnish-language education, media, and cultural production in the latter half of the nineteenth century. This historical pattern is consistent with previous studies that analyzed the Fennica metadata (Tolonen et al., 2016). However, a thorough analysis would require close collaboration with curators and literary researchers to evaluate the impact of cataloguing and institutional biases and gaps in data collections (Lahti et al., 2019).

Figure 2

Fennica metadata analysis (1809–1917). Top row: Publication volume trends by decade and annually. Middle row: Language distribution (total counts, temporal shifts, proportional changes). Bottom row: Top publication formats.

(Figure 3) examines metadata structure and language–format relationships in the Fennica corpus (1809–1917). Panel (a) shows the share of records per decade in which selected metadata fields are present. Language information is consistently complete across the period, while author coverage stabilizes at a relatively high level after early fluctuations. In contrast, subject coverage declines markedly over time, and series information, although initially rare, gradually increases toward the late nineteenth century. Panel (b) presents the corresponding missingness rates, highlighting the inverse pattern: subject and series fields exhibit substantial and increasing gaps in the mid- to late nineteenth century, whereas language remains fully populated and author missingness remains comparatively moderate. Panel (c) displays the distribution of languages within major format categories (shares calculated within each format). Finnish dominates most textual formats (e.g., books, journal articles, sheet music), while Swedish retains a visible presence across several formats and Russian appears primarily in specific categories such as maps. Overall, the figure demonstrates that metadata completeness and descriptive practices vary systematically across fields and over time, and that language use is structured differently across bibliographic formats.

Figure 3

Fennica (1809–1917): Format interaction and metadata quality. (a) Share of records per decade with Language, Author, Subjects, and Series fields present; dashed lines indicate LOESS trends. Language information is consistently near-complete, while subject coverage declines markedly over the century and series coverage gradually increases. (b) Field-level missingness by decade, showing complementary patterns to panel A, with high and persistent absence of series information and increasing subject gaps in later decades. (c) Distribution of publication languages across major formats (lumped categories); cell values represent the share within each format. Finnish dominates most textual formats, while language composition varies across material types such as maps and sheet music.

(Figure 4) examines how institutional concentration and subject description of the books evolve over time. Early decades are marked by extreme institutional concentration, with the majority of records contributed by a small number of institutions, as indicated by both high top-five shares and elevated Herfindahl–Hirschman Index (HHI) values (Hirschman, 1964). This concentration declines steadily across the nineteenth century, reflecting sustained diversification rather than short-term fluctuation. In parallel, subject diversity increases substantially, particularly toward the late nineteenth century, indicating both expansion of the published corpus and broader adoption of subject indexing practices. Subject churn analysis shows low continuity among dominant subjects in early decades, followed by increasing stability over time, consistent with the consolidation of recurring thematic categories and the maturation of cataloguing workflows. Together, these trends point to a transition from a highly centralized and sparsely described bibliographic system to a more distributed and thematically structured national bibliography.

Figure 4

Fennica (1809–1917): Concentration and subject dynamics. (a) Share of records contributed by the five largest libraries per decade, showing a gradual decline in institutional concentration over time. (b) Herfindahl–Hirschman Index (HHI) of library shares, confirming decreasing concentration and subsequent stabilization at lower levels. (c) Number of unique subject labels per decade, indicating strong growth in thematic diversity toward the late nineteenth century. (d) Jaccard similarity between top-50 subject sets in consecutive decades, showing increasing stability of dominant subject categories. Black points and solid lines show observed values; blue dashed lines indicate LOESS smoothing.

Taken together, Figures 3 and 4 show that the development of the Fennica corpus between 1809 and 1917 is characterized by uneven and asynchronous development across metadata quality, language use, institutional participation, and thematic description. Improvements in metadata are field-specific rather than uniform, with persistent gaps in some descriptors and delayed adoption of others, particularly subject indexing. Language use evolves gradually and in close interaction with publication format, reflecting system-wide cultural change rather than isolated shifts within specific content type. At the same time, the corpus transitions from a highly centralized institutional structure to a more distributed one, accompanied by increasing subject diversity and thematic stabilization. These combined trends indicate that changes in coverage, description, and structure are driven by evolving cataloguing practices and institutional expansion rather than by simple linear growth, underscoring the need for temporally aware and metadata-sensitive analytical approaches when working with historical bibliographic data.

4.3 Case study 2: Finnish music metadata

Having examined bibliographic metadata in Fennica, we now turn to music heritage data in the Viola collection, which presents different metadata challenges and opportunities. As another example, let us look at patterns in Finnish music production. The package allows users to access records of Finnish music-related items such as books, music melody, harmony, rhythm, lyrics, recordings, and archival materials. We can use search terms and apply specific filters, for example, to retrieve score or audio formats or focus on particular media types, languages, time periods, or institutions, thus streamlining the process of gathering relevant cultural data for analysis (National Library of Finland, 2023). This supports the study of trends over time, analyzing the types of musical compositions archived, and examining cultural themes in Finnish music.

Using the finna package, we can retrieve the Viola collection. We created a shortcut function to systematically fetch the whole metadata records for the Viola collection by dividing them into years depending on the user needs and also the chance to divide the records into batches of years to avoid the bottleneck of 100,000 records per request (National Library of Finland, 2023). Similarly, this function also includes records with missing years, providing the user with complete metadata. We used the finna package to gather 1,149,655 metadata records from Viola. The collection was retrieved systematically by dividing records into manageable year-based batches to provide a complete picture of its content and structure.

The number of music-related records increased until the mid-1900s, and then after that the growth became much faster (Figure 5). There is a noticeable rise in records starting from the 1960s, with a sharp increase from the 1980s onward, especially during the early 2000s. This increase likely reflects intensifying digitization and systematic cataloguing efforts beginning in the late twentieth century. Looking at the languages used in the records, Finnish is the most common, followed by English. Many entries are also marked with a special code (zxx), which means “no linguistic content, inapplicable.” Swedish, German, and other languages also appear in smaller amounts. When we look at how language use has changed over time, we see that English has become common in recent decades, especially for international music and learning material.

Figure 5

Trends and language composition in the Viola collection. (a) Number of records by decade (1900–2025). (b) Annual record counts with a smoothed trend line. (c) Most common publication languages. (d) Changes in publication language counts across decades. (e) Relative language composition by decade across the full temporal coverage of the collection. (f) Proportional distribution of sound recordings and sheet music by decade.

The formats in the Viola collection are mostly audio recordings and sheet music. Other formats such as books, text documents, and videos are present as well, but in much smaller numbers. This shows us that Viola is not just a collection of written works but also includes a wide variety of musical materials.

Overall, these patterns provide a structural overview of how the publishing activities catalogued in Viola have changed over time across publication volume, language use, and material format. The Viola collection provides an underutilized resource for the study of Finland’s musical history and its preservation and dissemination over the years.

(Figure 6) examines post-1900 music metadata in the Viola collection by comparing the three dominant formats: sheet music, sound recordings, and video, and shows that language patterns are strongly format dependent. Panel (a) demonstrates the structural shift in production: sheet music grows steadily across the twentieth century, sound recordings expand dramatically in the late twentieth century and become numerically dominant, and video emerges later at a smaller scale. Panel (b) shows that Finnish overwhelmingly dominates sheet music, while sound recordings and especially video display greater linguistic diversity and a larger share of non-linguistic (instrumental) content. Panel (c) confirms in absolute terms that Finnish remains the largest language across formats, but English and non-linguistic content become increasingly significant in recordings and video. Overall, the figure indicates that language shifts after 1900 are not uniform across music metadata but closely tied to technological change.

Figure 6

Music metadata after 1900: format-specific language dynamics in the Viola collection. (a) Annual growth of sheet music, sound recordings, and video. (b) Overall language composition within each format. (c) Absolute counts of the three most frequent languages per format.

5 Discussion

We present finna, a research-oriented R interface that integrates REST and OAI-PMH access to the Finnish national cultural metadata hub Finna. The package automates large-scale metadata retrieval, manages API constraints, and transforms heterogeneous records into tidy and analyzable data structures within reproducible R workflows. This addresses a critical shortcoming in cultural heritage research infrastructure by facilitating statistical programming of metadata structures, institutional coverage, and longitudinal patterns in cultural production. The primary target audience for the work includes quantitatively oriented digital cultural heritage researchers and data scientists.

Our case studies illustrate both the capabilities of automated metadata retrieval and the analytical potential of large-scale bibliographic and music metadata. They demonstrate how automated metadata retrieval can be used to examine large-scale descriptive patterns in bibliographic and music-related metadata, rather than to directly explain historical processes. This work contributes to ongoing efforts in bibliographic data science by facilitating transparent access to nationally aggregated cultural metadata while making structural biases and quality issues explicitly traceable in reproducible workflows. Beyond data retrieval, the software also supports metadata refinement and exploratory visualization for quality control purposes. While full automation of metadata refinement is often infeasible due to the need for domain expertise and contextual interpretation, such methods can highlight potential structural biases and guide expert-driven refinement. Recent efforts in bibliographic data science have shown how large-scale analysis of cultural metadata can reveal historical transformations in public discourse, publishing practices, and knowledge production based on harmonized metadata and scalable open science workflows (Lahti et al., 2019, 2020; Tolonen et al., 2019a, b).

A closer examination of existing R and Python packages for accessing cultural heritage APIs, including those mentioned in the Introduction, suggests that their primary contribution is to facilitate basic programmatic access to individual data sources. Examples include the R packages rdpla and europeanaR, as well as Python tools such as pyeuropeana, DPyLA, and the Smithsonian Open Access API client. While these tools are valuable for exploratory queries and targeted data retrieval, they do not support metadata refinement or the creation of analysis-ready datasets for exploratory and statistical analysis within R. In particular, existing packages typically leave responsibility for managing API constraints, pagination limits, metadata quality issues, and downstream integration to the user. We have aimed to address these shortcomings in the finna package.

The importance of metadata in cultural heritage research has led to the development of various tools and platforms to improve access to structured data from galleries, museums, libraries, and archives. Metadata is a key resource for digital humanities research, enabling scholars to explore patterns across historical periods (Fear, 2010). The evolution of semantic portals for cultural heritage research has progressed from data aggregation and harmonization and improved search and browsing functionalities to interactive tools and, more recently, to the integration of artificial intelligence for data processing, analysis, and interpretation (Hyvönen, 2020). Recent advances in natural language processing offer opportunities to explore and enrich metadata beyond descriptive statistics through named entity recognition and semantic similarity analysis of titles, subject headings, and descriptive fields. This could support enhanced exploratory analysis and hypothesis generation (Manning et al., 2014). Their application, however, requires careful attention to metadata quality, multilingualism, and historical bias, particularly in aggregated collections spanning long periods and across diverse institutions.

The work provides limited capabilities for automating tasks that require expertise and customization. The structured access and transparent workflows provided by the finna package are prerequisites for large-scale statistical analysis. Our work demonstrates the approaches to metadata retrieval, refinement, and exploratory quality control. The software will be maintained and extended to support research. In Finland, the development of the Finna platform and its API was driven by the need to provide unified access to cultural heritage metadata from a wide range of institutions, including national libraries, museums, and archives. Concerns related to bias, provenance, and the responsible reuse of cultural heritage data have been further formalized in emerging best-practice frameworks, such as the CARE principles (Collective Benefit, Authority to Control, Responsibility, and Ethics) (Carroll et al., 2020). Complementing this perspective, recent work on datasheets for digital cultural heritage datasets proposes structured documentation practices to make dataset provenance, selection processes, limitations, and potential biases explicit (Alkemade et al., 2023). Initiatives such as the European data space for cultural heritage aim to support large-scale reuse of cultural heritage data. Within this landscape, our contribution is a researcher-facing tool that enables reproducible access to national-level cultural metadata while remaining compatible with broader European efforts toward harmonized cultural heritage data spaces. Looking ahead, structured and reproducible access to cultural metadata provides a foundation for more advanced analytical approaches.

Several opportunities could be foreseen to expand the functionality of the finna package. One significant improvement is the development of a Shiny web application that integrates with the finna R package. This would provide an intuitive graphical user interface (GUI) for querying metadata from the Finna-API, making Finnish cultural heritage data more accessible to researchers and the general public. Improved visualization tools could be incorporated to allow users to create interactive plots and visual summaries of metadata. This will support analyses of Finnish cultural heritage and its change over time.

Conclusion

This study promotes the use of computational and automated workflows as an emerging approach for studying cultural heritage using structured, tabular metadata collections. We demonstrate this approach based on the finna R library to retrieve, process, and analyze Finnish cultural metadata. By facilitating programmatic access, the package enables researchers to integrate collections of cultural heritage data into reproducible computational workflows within R. Through case studies on Fennica and Viola, we demonstrate the package’s capacity to support cultural analytics, illustrate temporal trends, and facilitate large-scale metadata research. The work addresses the gap in digital humanities infrastructure by connecting high-quality metadata resources with widely adopted analytical tools, making cultural data more accessible and reusable. It also supports the principles of open science by promoting transparency, reproducibility, and modularity towards novel extensions. As digital archives continue to expand, open-source tools such as finna can support data-driven engagement with cultural heritage collections across research, education, and citizen science.

Author Contributions

All authors read and approved the final version of the manuscript. L.L conceptualized the software; A.J. and L.L. created the package; A.J., J.M., and L.L. authored and reviewed the manuscript.

Language: English
Page range: 27 - 27
Submitted on: Oct 15, 2025
Accepted on: Jul 6, 2026
Published on: Jul 24, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Akewak Jeba, Julia Matveeva, Leo Lahti, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.