Skip to main content
Have a personal or library account? Click to login
Semi‑Automatic Pipeline for the Transcription of Mensural Polyphony into Symbolic Interpreted Scores Cover

Semi‑Automatic Pipeline for the Transcription of Mensural Polyphony into Symbolic Interpreted Scores

Open Access
|Jul 2026

Full Article

1 Introduction

This article presents the first system for the semi‑automatic transcription of Late Medieval and Renaissance polyphonic sources in mensural notation into fully interpreted symbolic scores encoded in Mensural MEI.1 The system integrates multiple early music technologies to address the specific challenges posed by this repertoire— most notably, the rhythmic interpretation of mensural notation, which fundamentally distinguishes it from Common Western Music Notation (CWMN). The resulting workflow produces vertically aligned, interpreted scores suitable for both musicological research and performance by modern musicians.

This work is part of a larger project aimed at making the digitization and encoding of mensural music sources as accessible and efficient as possible, enabling global adoption, including by institutions with limited resources. The proposed workflow was applied to a handwritten polyphonic choirbook (GCA‑Gaha 1), held at an institution without the infrastructure or resources required to digitize such material (i.e., over‑sized bound volumes belonging to a special collection). This case study therefore served as a pilot project to assess the practical applicability of the proposed digitization and music information retrieval (MIR) pipeline in a low‑resource institutional context (see Figure 1).

Figure 1

Digitization and music information retrieval (MIR) pipeline showing the input and output of each step. The MIR part (green boxes) allows for the semi‑automatic transcription of mensural sources into interpreted scores.

The details of the digitization phase of the workflow (orange box in Figure 1), which involved the use of a do‑it‑yourself (DIY) book scanner and the adaptation of digitization guidelines for institutions with limited resources, are described in Thomae et al. (2022b). The present paper focuses on the MIR components of the workflow (green boxes in Figure 1). In line with the project’s accessibility goals, the technologies employed in the MIR pipeline must (1) be free and open access, (2) support a high degree of automation to reduce human labor, and (3) provide interfaces that facilitate necessary manual correction of automatic outputs. Rather than developing a new monolithic application, this project adopts an approach based on the reuse, enhancement, and interoperability of existing tools, each specialized in a particular MIR task. Specifically, the system integrates MuRET for the automatic recognition of music symbols, the Measuring Polyphony (MP) Editor for the automatic rhythmic interpretation (or ‘scoring‑up’) of mensural notation, and the Humdrum dissonance analyzer to detect and assist in correcting potential scribal or recognition errors that result in vertical misalignments.

The system addresses a long‑standing gap in optical music recognition (OMR) workflows by enabling the rhythmic interpretation of mensural notation, which is essential for transcribing mensural parts into symbolic scores. These machine‑readable scores can increase access to this repertoire for the general public (through playback) and for modern musicians (through transcription into modern values). Beyond access, these scores are of interest to early‑music scholars: by encoding both melodic and vertical intervallic information within a unified representation, the symbolic scores facilitate historical and digital musicology research, allowing scholars to evaluate contrapuntal structures and voice‑leading patterns, and to conduct cross‑source comparisons and large‑scale corpus analyses through computational methods. These capabilities extend those of existing OMR systems significantly.

2 Mensural Notation

Mensural notation was the main system used for writing polyphonic music during the Late Middle Ages and Renaissance. Most surviving sources employing this notation present the music in a separate‑parts layout, where each voice is copied in a different part of the page or opening (see Figure 2) or in separate partbooks. While this layout fully conveys the melodic content of each voice, it obscures the vertical sonorities—that is, the simultaneous sounding pitches across voices. These vertical relationships become apparent only when the voices are either sung together or transcribed into a modern score format, where all voices are vertically aligned. However, such a transcription (or performance) requires specialized knowledge, as mensural notation encodes rhythmic values differently from CWMN.

Figure 2

Example of voices written in choirbook layout; each voice is in a different quadrant of the book opening (Thomae et al., 2022a).

In mensural notation, a note’s duration is determined not solely by its shape but also by several other factors. The first of these is mensuration, which is broadly analogous to the modern concept of meter. Mensuration operates at the following hierarchical levels:

  • Modus: Relationship between longa (tismir-9-1-292-ug7.png) and brevis (tismir-9-1-292-ug8.png).

  • Tempus: Relationship between brevis (tismir-9-1-292-ug8.png) and semibrevis (tismir-9-1-292-ug9.png).

  • Prolatio: Relationship between semibrevis (tismir-9-1-292-ug9.png) and minima (tismir-9-1-292-ug10.png).

Each of these levels may consist of three subdivisions (called perfect or major, depending on the level) or of two subdivisions (imperfect or minor) (see Table 1). For example, in a piece written in perfect modus, imperfect tempus, major prolatio, longas are equivalent to three breves, with each breve being equivalent to two semibreves and each semibreve equivalent to three minimas by default. These default values form the structural basis for rhythmic interpretation.

Table 1

Mensuration values.

MODUS
longa–brevis relation
tismir-9-1-292-ug1.pngtismir-9-1-292-ug2.png
perfect modusimperfect modus
perfect longa by defaultimperfect longa
TEMPUS
brevis–semibrevis relation
tismir-9-1-292-ug3.pngtismir-9-1-292-ug4.png
perfect tempusimperfect tempus
perfect breve by defaultimperfect breve
PROLATIO
semibrevis–minima relation
tismir-9-1-292-ug5.pngtismir-9-1-292-ug6.png
major prolatiominor prolatio
perfect semibreve by defaultimperfect semibreve

In perfect mensurations (i.e., those based on ternary divisions), the actual duration of a note can be influenced by its context—that is, the rhythmic environment created by surrounding notes. Contextual interpretation can override the default note values through two primary mechanisms:

  • Imperfection: A note’s value is reduced from perfect (three units) to imperfect (two units) (see Figure 3a).

  • Alteration: A note’s value is doubled (see Figure 3b).

Figure 3

Example of imperfection (a) and alteration (b). Both examples (a and b) show the original in mensural notation on top and its modern transcription below. (a) Same note shape with a perfect (P; triple) and imperfect (I; duple) value. (b) Same note shape with a regular (r; one beat) and altered (A; two beats) value. Notice the presence of a dot of division between the first and second semibrevis.

The principles of imperfection and alteration, summarized by the theorist Franco of Cologne in his treatise Ars cantus mensurabilis (ca. 1280), describe the contexts in which these two modifications apply. These principles can be consulted in the Strunk and Treitler (1998) translation and the books by Apel (1953), Parrish (1978), and DeFord (2015). In addition to contextual factors, certain graphical features can also affect note values (Kelly, 2015; Rastall, 2010):

  • Coloration: Use of a different ink color (or filled‑in note‑heads) to indicate a change in rhythmic interpretation. Its effect depends on the prevailing mensuration.

  • Dots of division (punctum divisionis): Clarify rhythmic groupings in triple meter by marking the division of notes into perfect (ternary) units (see Figure 3b). They function similarly to barlines, organizing notes into distinct rhythmic groups.

  • Dots of addition (punctum additionis / augmentationis): Increase a note’s value by half, akin to the modern dotted note. Visually identical to dots of division, their meaning must be inferred from context.

Given the notational complexity and the separate‑parts layout of the sources, producing vertically aligned, interpreted MEI scores demands both rhythmic understanding and symbolic encoding. Moreover, assessing the correctness of this alignment can benefit from consideration of the piece’s counterpoint. The following sections examine OMR, mensural notation interpretation, and counterpoint analysis tools that can be used to address these challenges.

3 Technologies for Late Medieval and Renaissance Polyphonic Music and Its Notation

3.1 OMR frameworks for mensural notation: MuRET

OMR converts images of music documents into machine‑readable symbolic formats (e.g., MusicXML, MIDI, MEI, Humdrum). For early music, the Music Encoding Initiative or MEI format is often preferred because it supports neumes and mensural notation (Hankinson et al., 2011; Roland et al., 2014). MEI also links encoded notes to image regions via its facsimile module.

There are two OMR systems that support all stages of mensural notation recognition with user‑friendly interfaces: Aruspix and MuRET. In addition, the Single Interface for Music Score Searching and Analysis (SIMSSA) project (Fujinaga et al., 2014) employs an OMR workflow primarily used on neumatic notation (especially square notation) but incorporating machine‑learning components that can be adapted to other sources and notation types.

Regarding the SIMSSA OMR workflow, although most of its components can be applied to mensural notation—such as the document analysis and symbol classification stages—the components involved in the final stage of the workflow—music encoding (via the Music Encoding job) and its correction (via the Neume Editor Online, Neon (Regimbal et al., 2020))—are specifically designed for neumatic notation. While adapting the Music Encoding job to support mensural notation may be relatively straightforward, extending the full workflow to this notation would still require the development of a dedicated online mensural editor to replace Neon.

Aruspix, a macOS desktop tool developed by Pugin (2018) in collaboration with the Marenzio Online Digital Edition and the Music Encoding Initiative, enables OMR, superimposition, and collation of early typographic prints. However, because it is limited to printed mensural notation and is no longer supported on recent macOS systems, it could not be applied to the handwritten source GCA‑Gaha 1.

The MuRET (‘Music Recognition Encoding and Transcription’) framework, developed by David Rizo within the Hispamus project,2 supports OMR of both printed and handwritten sources in mensural and CWMN (Rizo et al., 2018a).3 MuRET has been used to encode mensural sources from Zaragoza (Rizo et al., 2020) and the National Library of Spain (Rizo et al., 2022). Because of this, it already includes models that have proven effective on handwritten Spanish sources, which share visual and notational characteristics with the GCA‑Gaha 1 choirbook, such as the more droplet‑like note heads common in Iberian manuscripts. These models provide a strong starting point for training new ones specifically adapted to the Guatemalan source.

While substantial research has been conducted on OMR for mensural notation—see Pacha and Calvo‑ Zaragoza (2018) and Ríos Vila et al. (2022)—this project focuses on systems that not only support the full OMR process but also provide user‑oriented interfaces for correcting the OMR output.

3.2 Rhythmic interpretation tools for mensural notation: Scoring up

The separate‑parts layout of Late Medieval and Renaissance polyphonic sources in mensural notation obscures the visualization of the vertical sonorities. It is not until the music is transcribed into a score or musicians sing the parts together that the polyphonic texture of the piece can be appreciated (Owens, 1997; Cumming, 2013). In a score layout, the reader has access to the same melodic information as in the original separate‑parts layout, plus the information about the vertical intervals. However, lining up the mensural parts into a score requires the interpretation of the duration of their notes, which implies dealing with the complexities of the notation: the mensuration, note shapes, coloration, dots (of addition or division), and—in triple meter—with the context‑dependent nature of rhythm. Moreover, potential scribal errors must also be considered.

Researchers have developed a few tools to deal with the rhythmic interpretation of mensural notation. As part of their OMR research, Huang et al. (2015) presented a framework that involved the transcription of mensural notation into modern values; however, the transcription work is restricted to imperfect mensuration, which constitutes the trivial case of the problem. On the other hand, Rizo et al. (2017) presented a mensural‑to‑modern state transducer as part of a larger interactive editing system to assist musicologists in transcribing XVI–XVIII c. Hispanic polyphony. The idea behind this transducer was for it to learn about common scribal errors based on the musicologist corrections of the initial interpretation of the notes; however, there has been no report on the system’s performance.

Finally, Thomae et al. (2019) introduced an automatic voice alignment (or ‘scoring‑up’) tool for mensural music, an expert system for the rhythmic interpretation of mensural notation that handles the context‑dependent nature of mensural music in triple meter—the nontrivial case—based on the principles of imperfection and alteration. The mensural scoring‑up tool is a Python command‑line application that takes as input Mensural MEI files encoding each part (i.e., voice) of the piece and outputs a single Mensural MEI file that preserves the original information and adds the interpretation of each note as perfect, imperfect, or altered, as well as the interpretation of dots as dots of division or addition. When the resulting MEI file is provided to Verovio—the established MEI rendering library—the piece is displayed as a score with all the voices lined up.4 The tool performed well on a XIV–XV c. corpus, with the main source of error being missing or misplaced dots of division in the original sources (Thomae et al., 2019). Given these results, the mensural scoring‑up tool was selected to generate a ‘draft scored‑up’ version of the pieces in GCA‑Gaha 1, which could then be corrected for scribal errors by an expert.

Later, Plaksin and Lewis (2022) developed the Mensural Rhythm Interpreter Tool (MeRIT), a client‑side JavaScript tool that interprets durations in mensural notation using rules derived from the writings of 15th c. music theorist Johannes Tinctoris.

3.3 Computational analysis tools, counterpoint, and renaissance polyphony: Humdrum

There are a few computer‑aided musicology tools that can be used to analyze counterpoint in Renaissance polyphonic music. These include the Python music analysis libraries of VIS Framework5 and CRIM Intervals,6 both built on top of the well‑established computer‑aided musicology toolkit of music21,7 and the Humdrum analysis tools, which were used in this study and are the focus of this section.

Humdrum is a system for computational musicological research created by David Huron in the 1980s and currently maintained by Craig Sapp at the Center for Computational Assisted Research in the Humanities at Stanford University. It consists of two main components: a symbolic encoding format (or Humdrum syntax) and a suite of software for analyzing the encoded data. The Humdrum format for CWMN is known as Humdrum **kern, while Humdrum **mens is used for mensural notation (Rizo et al., 2018b). The Humdrum analysis software includes the following tools:

  • Huron’s original Humdrum Toolkit, a set of Unix command‑line tools (written in Bash and AWK) to parse and analyze Humdrum data.8

  • The Humdrum Extras, a set of command‑line tools (in C++) that expand the capabilities of the Humdrum toolkit.9

  • Humlib, a C++ parsing library for Humdrum data files.10

  • The Verovio Humdrum Viewer (VHV) editor, an online Humdrum file editor with integrated graphical notation display through Verovio (supported by VHV’s internal conversion of Humdrum data into MEI).11 Humlib functionality is available within VHV through its compilation into JavaScript.

As indicated before, all these analysis tools work on Humdrum data; therefore, formats like MusicXML and MEI must be converted to Humdrum **kern first, which is possible using humlib’s musicxml2hum and mei2hum commands. In the VHV Editor, users can directly load MusicXML and MEI files, as these are internally converted to Humdrum to allow for the use of humlib’s processing tools, which are represented as filters in the VHV Editor. Filters are commands that modify Humdrum data. While most of them are related to data processing (e.g., filters to extract a part from a score), more advanced filters are available and can be accessed through the Analysis menu of the VHV Editor (see Figure 4).12 The user can use these to analyze Renaissance music, exploring various aspects such as imitation, melismas, and dissonances. The latter involves the use of the dissonance filter (DF) developed by Morgan (2017) to identify and index the various types of dissonances present in a Renaissance piece (e.g., passing tones, neighbor tones, suspensions).

Figure 4

Verovio Humdrum Viewer with the encoding of the piece in Humdrum **kern (left) and its Verovio rendering (right).

4 Transcription Pipeline of Mensural Sources into Interpreted Scores

This section introduces the first full MIR pipeline for transcribing mensural music sources into interpreted scores. These interpreted scores are symbolic files in which the original mensural note durations are decoded and the individual voice parts—originally notated separately—are aligned in a modern score layout. This alignment makes both melodic and vertical intervallic relationships visible at any moment in the music. As shown in Figure 1, the pipeline brings together three existing technologies: the MuRET optical music recognition framework, the mensural scoring‑up tool, and humlib’s DF—with the last two integrated within MP Editor (see Section 4.2). The scope of these tools is shown in Table 2.

Table 2

Scope of the tools in the transcription pipeline.

MURETMP EDITOR
SCORING‑UP FUNCTIONALITYDISSONANCE FILTER (DF)
Recognition of black and white mensural symbolsArs antiqua, Ars nova, and early Renaissance (does not work well on triple meter pieces written toward the end of the Renaissance, when alteration starts falling into disuse)Designed for Renaissance music

The following sections provide details on the tools involved in the pipeline, the work carried out to ensure their integration and interoperability, and the resulting symbolic corpus: the GCA‑Gaha 1 Mensural MEI scores.

4.1 MuRET

The MuRET framework guides the user through the OMR workflow via four main interfaces:

  1. Document analysis: Uses a selectional auto‑encoder to detect staff regions (Castellanos et al., 2020), with options for manual correction.

  2. Parts: Enables manual multi‑staff selection and assignment of staves to musical parts (e.g., alto, tenor).

  3. Transcription: Generates both agnostic and semantic representations of the music contained in each staff region.

    1. The agnostic representation consists of a sequence of tokens encoding only the graphical information of each symbol (specifically, its position in the staff—the line/space number—and its symbol class). It is obtained with a holistic staff‑level recognition algorithm for mensural notation based on a recursive convolutional neural network (Calvo‑Zaragoza et al., 2019). The recognized symbols can be corrected by modifying them using the interface tools to change their position in the staff or their category (see Figure 5a), by deleting them, or by adding a new symbol using a symbol‑level classifier available in the interface (Iñesta et al., 2019).

    2. The semantic representation, based on Humdrum **mens (Humdrum encoding for mensural notation) (Rizo et al., 2018b), is derived heuristically from the agnostic representation via a transducer (Rizo et al., 2017). This semantic encoding can also be corrected through the interface by modifying the **mens encoding of the notes (see Figure 5b).

  4. Document overview: Displays document pages, allowing page selection for MEI export.

Figure 5

MuRET’s transcription interface with agnostic and semantic transcription panels and correction options. (a) Agnostic transcription of the staff‑region image, with correction options for the symbol class (left) and for the line/space positions within the staff (top‑right from agnostic panel). (b) Semantic transcription encoded in **mens and rendered by Verovio, with correction option through the **mens table (left).

4.2 MP Editor

MP Editor is an online editor for mensural notation,13 developed as part of the Measuring Polyphony Project, led by Karen Desmond.14 It has various steps to enter a mensural piece:

  1. Upload page. In this initial step, users either (1) provide a URL linking to the source of the piece to be transcribed or (2) upload a parts‑based Mensural MEI file that follows MP Editor’s encoding conventions. For option (1), the Editor currently supports URLs from Gallica and eCodices, as the sources must include a IIIF manifest to display folio images within the Editor. For option (2), the uploaded file is typically one previously created using MP Editor, which the user wishes to reopen to continue editing.

  2. Metadata‑entry page. Here, users provide basic metadata about the piece to be transcribed, including title, composer, notation (with three options available: black mensural Ars antiqua, black mensural Ars nova, and white mensural), siglum, genre, and contributor’s name and role (encoder, proofreader, or editor).

  3. Input editor. Here, users select each staff region of the piece and enter its music and text using the computer keyboard. The information entered for each staff is internally stored in Humdrum **mens. Additionally, users can specify the part (i.e., voice) to which the selected staff belongs via a drop‑down menu and set the mensuration using radio buttons (see Figure 6a). Once the encoding is complete, users can download a parts‑based MEI file representing the full piece as a collection of parts (without rhythmic interpretation).

  4. Score editor. Here, note durations are interpreted and the voices are lined up automatically into a score, thanks to the integration of the mensural scoring‑up tool into MP Editor. Users can download two types of MEI files encoding the full piece: the parts‑based MEI file encoding the piece as a set of parts (without rhythmic interpretation) or a score‑based MEI file encoding the piece as a score with the voices lined up according to the notes’ duration (see the Supplementary Material for the structure of both file types). Additionally, users can continue correcting the piece within the score editor and save these changes as editorial, provided the ‘Continue in Editorial Mode’ option is selected (Figure 6b). In this case, the score‑based MEI will store both the original and corrected readings.

Figure 6

Measuring Polyphony Editor. (a) Input editor. (b) Score editor.

4.3 Integration of the scoring up into MP Editor

As part of the Measuring Polyphony Project, I incorporated the mensural scoring‑up tool into MP Editor. This implied rewriting this Python command‑line tool into JavaScript to work within the MP score editor, and adding some functionalities to the scoring‑up tool and the MP input editor:

  • Support for the scoring‑up tool to deal with the older notation style of black mensural Ars antiqua.

  • Support for the scoring‑up tool to handle changes in mensuration.

  • Support for entering colored notes in the input editor. This, at the same time, implied adding support in **mens for encoding colored notes (details in Section 4.5).

  • Support for entering separately the information about the mensuration sign and its meaning in the input editor. The mensuration sign is assigned through a drop‑down menu, while the mensuration values are assigned through the radio buttons for modus, tempus, and prolatio (see Figure 6a). This implied adding a way in **mens for encoding both the mensuration sign and its semantics separately (see details in Section 4.5).

4.4 Interoperability between MuRET and MP Editor

The integration of MuRET and MP Editor aims to streamline the transcription workflow by leveraging the automatic features of both tools, thereby reducing the need for manual input. MuRET’s OMR capabilities eliminate the need to manually enter musical notes into the MP input editor, as these are automatically extracted during the OMR process. Conversely, MP Editor’s automatic scoring‑up functionality computes the perfect, imperfect, and altered rhythmic values of the notes, removing the need to manually assign them in MuRET.15

To enable MuRET’s symbolic output to be used as input in MP Editor, its Mensural MEI file must be formatted in a way that is compatible with MP Editor’s upload system. This system is designed to accept parts‑based Mensural MEI files, typically exported from MP Editor itself, allowing users to save and resume their work seamlessly. As a result, the first requirement was for MuRET to support exporting a parts‑based Mensural MEI file.

This exported file follows a set of conventions concerning the following aspects: the inclusion of references to the IIIF manifest of the source and links to each of the images involved via their IIIF URIs; the encoding of structural elements such as page and system beginnings; and the encoding of notational features such as mensuration signs, clefs, and key signatures. For full details on these conventions, see Desmond et al. (2021).

4.5 Interoperability between MP Editor and the humlib DF

Like how MuRET allows users to verify and refine the output of the OMR process, MP Editor enables users to review and correct the results of the automatic scoring‑up process through its score editor. Misalignments in the resulting score can stem from undetected OMR errors, scribal mistakes in the original source, or inaccuracies in the scoring‑up algorithm.

MP Editor already supports error detection through three features:

  1. Displaying the piece in score format, which reveals vertical sonorities that are not apparent in the original, part‑based layout.

  2. Allowing flexible barring based on different note values.

  3. Supporting the use of modern clefs to improve readability.

These capabilities facilitate the analysis of counterpoint and help users identify inconsistencies more effectively.

Building on this foundation, the integration of humlib’s DF enhances the MP Editor by explicitly highlighting counterpoint violations. This integration aimed to evaluate whether drawing attention to musically implausible passages would aid in identifying and resolving scribal errors—an expectation that proved accurate (Thomae et al., 2022a).

Humlib was chosen over other counterpoint analysis libraries, such as VIS and CRIM Intervals, for several reasons. Most notably, Humdrum **mens is already integrated in both MuRET and MP Editor. In addition, humlib’s DF supports a broader range of dissonance types than the VIS Framework’s dissonance indexer, allowing for more precise classification. At the outset of this project, CRIM Intervals also lacked functionality for dissonance detection.

Because humlib’s DF operates exclusively on **kern data (Humdrum’s format for CWMN), a conversion workflow is required to transform MP Editor’s score‑based Mensural MEI files into **kern format to apply the DF and retrieve the dissonance labels for display within MP Editor. This conversion workflow, previously introduced in Desmond et al. (2021), can be summarized as follows:

  1. MP Editor’s score‑based Mensural MEI output is first converted into **mens format, while preserving the XML IDs of MEI notes and rests.

  2. The **mens file is then converted into **kern format, again maintaining the original MEI XML IDs throughout.

  3. The humlib DF is applied to the **kern file. At this stage, the output includes three spines per voice: a **kern spine (musical events), a **xmlid spine (MEI identifiers), and a **cdata spine (dissonance labels); see the example in Table 3 corresponding to the piece in Figure 7.

  4. Dissonance labels from the **cdata spine are extracted, along with their associated note IDs in the **xmlid spine, and compiled into a JSON list that pairs each note ID with its corresponding dissonance classification (see Figure 8).

  5. These dissonance labels are displayed in MP Editor as lyrics attached to the corresponding notes (matched via XML ID). The labels are color‑coded: blue for dissonances with explainable functions (e.g., passing tones, neighbor tones, suspensions) and orange for dissonances that cannot be explained according to Renaissance counterpoint rules. For more on these categories and their criteria, see Thomae et al. (2022a).

Table 3

Example of the conversion from **mens to **kern and the application of the DF on a three‑voice piece. Each group of three adjacent spines provides information for a different voice, tenor (red), altus (green), and superius (blue). The **kern spine contains the voice’s notes in CWMN values, the **xmlid spine contains the corresponding XML IDs of these notes as encoded in the original MEI file, and the **cdata spine contains their corresponding dissonance labels.

!!!OTL: 6 Missa sobre las voces
!!!system‑decoration: [(s1,s2,s3)]
**kern**xmlid**cdata**kern**xmlid**cdata**kern**xmlid**cdata
*part3*part3*part3*part2*part2**part1*part1*
*staff3*staff3*staff3*staff2*staff2**staff1*staff1*
*I"tenor***I"altus***I"superius**
*clefC3***clefC2***clefG2**
met()_0222***met()_0222***met()_0222**
*MM600***MM600***MM600**
1GPART3_A12.00r..00r..
2GPART3_A15.......
2APART3_A16.......
===.=====
2BPART3_A17.......
2cPART3_A18.......
1dPART3_A19.......
=========
1.ePART3_A21.00r..1gPART0_A20.
......2gPART0_A23.
4dPART3_A24....2aPART0_A25.
4cPART3_A26.......
===.=====
2dPART3_A28....2bPART0_A27.
2ePART3_A30....2ccPART0_A29.
2fPART3_A31....1ddPART0_A32.
2.gPART3_A33.......
=========
...1cPART6_A35.2eePART0_A34.
4fPART3_A36p......
4ePART3_A37....1ggPART0_A38.
4dPART3_A39n......
2ePART3_A40.2cPART6_A41....
2dPART3_A43.2dPART6_A44.2ffPART0_A42.
=========
2cPART3_A46.2ePART6_A47.2eePART0_A45.
2dPART3_A48.2fPART6_A49.1ddPART0_A50.
1BPART3_A51.1gPART6_A52....
......4ccPART0_A53v
......4bPART0_A54.
=========
0APART3_A56.0aPART6_A57.2ccPART0_A55.
=========
......2eePART0_A58.
......2eePART0_A59.
......2eePART0_A60.
=========
1r..2r..2aPART0_A61.
...2cPART6_A64.1eePART0_A65.
2r..2cPART6_A67....
2ePART3_A69.2dPART6_A70V2ddPART0_A68V
=========
2ePART3_A72.2.ePART6_A73.2ccPART0_A71.
2ePART3_A75....4bPART0_A74.
...4fPART6_A77P4aPART0_A76n
2APART3_A79Z1gPART6_A80.2bPART0_A78Z
1dPART3_A82z...2eePART0_A81Z
=========
Figure 7

Rendering of the **kern encoded piece described in Table 3 in Verovio Humdrum Viewer. Ties were added to facilitate readability for notes that go across barlines; these ties are not present in the original file shown in Table 3.

Figure 8

JSON list showing the mapping of XML IDs and corresponding dissonance labels retrieved from the **kern file shown in Table 3.

The work done on integrating the DF and the MP Editor led to enhancements in mensural notation encoding and related software, including:

  • Enhancements in the **mens format to encode features already supported in Mensural MEI but not available in **mens yet. These are: rhythmic alterations (with the ‘+’ character), semantics of a mensuration sign (with the ‘met()’ encoding of the sign followed by an underscore and four numbers indicating the mensuration at the four different note levels), distinction between dots of division (with the new ‘:’ sign) and dots of addition (with the usual ‘.’ sign), and coloration (with the ‘’ character).16

  • A fix for issues related to the interpretation of perfect and imperfect durations in **mens and their conversion into Mensural MEI.

  • Expansion of humlib’s mei2hum conversion tool to cover, in addition to CMN MEI to **kern conversion, the Mensural MEI to **mens transformation.17

  • Creation of the humlib mens2kern conversion tool.

  • Improvement of MEI and Verovio support for encoding and rendering perfect, imperfect, and altered values.

Moreover, the improvements in **mens (the first two items) resulted in improvements to MuRET’s and MP Editor’s support for mensural notation, since **mens is used in MuRET’s semantic transcription and in the MP input editor to store the music entered by the user for each system.

4.6 The corpus of symbolic scores

This transcription pipeline was applied to the GCA‑Gaha 1 manuscript. The parts‑based and score‑based MEI files produced by MuRET and MP Editor can be consulted on GitHub.18 The score‑based MEI files (i.e., the transcribed scores) were subsequently modified to work with mei‑friend, an MEI editor that renders encoded music via Verovio and displays the facsimile (i.e., digital images).19 Most of these modifications were required to enable facsimile display in mei‑friend, which requires specific facsimile information to be provided within the MEI file, including references to the IIIF manifest and specific formatting of IIIF URIs. The final corpus of symbolic scores for GCA‑Gaha 1 is available in Zenodo and GitHub.20 Uploading the files to mei‑friend allows users to view the transcribed score with editorial corrections, play it back, and consult the original images simultaneously.

5 Discussion

This section evaluates the performance and limitations of the full transcription pipeline, with a focus on how each component contributes to or hinders the overall goal of reducing human intervention in the transcription of mensural polyphony.

5.1 Performance of MuRET

While the Document Analysis and Parts interfaces in MuRET performed efficiently,21 the most time‑consuming portion of the workflow occurred during the correction of the agnostic and semantic transcriptions within the Transcription interface.

5.1.1 Transcription: Agnostic representation

Issues encountered in the agnostic representation fall into three main categories:

  • Missing or underrepresented symbols in the training data of the holistic staff‑level recognition model. Although the training data included several annotated pages, the model consistently misidentified certain F clefs and failed to recognize all C1 clefs (C clef on the first line). This was due to gaps in the training set, which included only one style of F clef of the two used in the manuscript (mensural and modern) and lacked C1 clefs entirely. Beamed notes (tismir-9-1-292-ug11.png) were also commonly misclassified as fusas (tismir-9-1-292-ug12.png), reflecting their relative scarcity in the training data.

  • Failure to identify the custos. Despite being well represented in the training data, the custos was consistently missed by the holistic staff‑level recognition model. This may be due to its location at the extreme right of the staff, suggesting issues with the staff’s bounding box width—which might need to be extended sufficiently beyond the last symbol—or the image degradation along the manuscript’s edges. Attempts to add it manually using the symbol‑level classifier were unsuccessful. Although this symbol‑ level classifier was trained on a manuscript from Zaragoza—which shares many note shapes with the GCA‑Gaha 1 manuscript due to their common use of Hispanic mensural notation (characterized by droplet‑like note heads, as seen in Figure 2, rather than angular note heads)—the custos shape differs notably between the two sources. This discrepancy led to frequent misclassifications. While retraining an entirely new classifier tailored to the GCA‑Gaha 1 manuscript would involve significant data preparation and may be unnecessary given the overall similarity in notation, a more practical approach would be to fine‑tune the existing model using only those symbols—such as the custos—that diverge significantly between the two manuscripts.

  • Stacked symbols (symbols vertically aligned with others). Elements such as fermatas and accidentals above or below notes (or other accidentals in the case of multi‑accidental key signatures) were not handled properly by the holistic model, which was designed for monophonic staves and thus lacks support for vertical symbol alignment.

The first two issues could be addressed by enriching the training datasets of both recognition models. However, the third reflects an inherent design assumption and, therefore, requires manual correction under the current framework.

5.1.2 Transcription: Semantic representation

The most time‑consuming aspects of the semantic transcription involved correcting specific musical elements by manually editing the **mens table, which encodes the semantic representation of the score. These elements included the following:

  • Ligatures. All ligatures, regardless of their shape and their notes, were represented with a generic agnostic class ligature. As a result, their semantic encoding must be completed manually, entering each note in the ligature into the **mens table. MuRET’s new ‘Ligature’ panel for two‑note ligatures, which allows shape selection, individual pitch input, and toggle of coloration and dot features, alleviates some of this burden. However, extending this feature to support longer ligatures—assuming breves as default intermediate notes (except for ligatures cum opposita proprietate)—would further streamline this task. Recent updates allowing multiple ligature classes in the agnostic encoding may eventually eliminate the need for manual intervention.

  • Accidentals on the first note of a system. These were often misinterpreted as part of the key signature. Correction required deleting the key signature, adding the accidental manually to the **mens code of the corresponding note, and linking it to its bounding box in the image. Enhancing the agnostic‑to‑semantic conversion to check whether initial accidentals match accidentals typically found in mensural key signatures (B and E flat) could reduce this issue.

  • Accidental cancellation (gestural information). Gestural data—when the written accidental (encoded in the notes using the attribute accid) differs from its interpretation (encoded using the attribute accid.ges to represent ‘gestural’ or ‘performed interpretation’)—need to be added manually in the MEI file to reflect cancellation. A recent MuRET update now allows these data to be entered directly in the **mens table. Further automation could infer cancellation if an accidental shares the pitch position of a previous accidental within the same system.

  • Stem direction. Agnostic stem directions were sometimes overridden in the semantic encoding due to default CWMN rules (i.e., stems go down above the third line and up below it). This minor issue needs to be addressed in the agnostic‑to‑semantic conversion.

  • Associating accidentals and dots with the correct note. When placed between two notes, accidentals and dots were sometimes linked incorrectly. Correction required unlinking and relinking the elements via several manual steps. Context‑aware rules could help reduce this issue.

In the absence of ligatures and complex accidental scenarios—and setting aside minor issues such as stem direction—the agnostic‑to‑semantic conversion was largely accurate. When the agnostic representation was carefully corrected, further semantic adjustments were generally unnecessary.

5.1.3 Batch processing

MuRET includes three main automatic processes: document analysis, agnostic transcription, and semantic transcription. Despite their automation, each step currently involves a series of manual actions. Document analysis requires navigating to the Document Analysis interface, clicking the corresponding button to perform the analysis, and correcting the results. For transcription, the user must switch to the Transcription interface, select each individual staff region (identified during the document analysis), click a button to generate the agnostic transcription, correct it, then click a second button to generate the semantic transcription (which is based on the agnostic one) and correct that too. This sequence must be performed separately for every staff region across all folios, resulting in a time‑intensive and repetitive workflow. To streamline this process, a batch‑processing feature will soon be introduced. This addition will enable users to apply all automatic steps simultaneously across multiple staves and folios, allowing them to focus primarily on correcting results rather than performing repetitive interactions.

5.1.4 Text

MuRET currently lacks support for text recognition, and, as a result, the exported MEI files omit lyrics—a significant limitation for a corpus intended for vocal music. Future updates will incorporate text recognition functionality, and subsequent releases of the encoded corpus will include lyrics to provide a more complete representation of the sources.22 The text‑recognition process may be simplified by supplying a transcript of the lyrics, which are easily determined since the GCA‑Gaha 1 manuscript is a book of Masses, all of which share the same liturgical texts.

5.2 Performance of MP Editor and the DF

As reported in Thomae et al. (2022a), the use of the DF in MP Editor halved correction time by narrowing the error search area to the passage preceding the first orange dissonance label. It also improved correction accuracy by allowing users to experiment with changes and immediately see their effect on the voice alignment and dissonance labels. Moreover, the presence of the dissonance labels led to the discovery of a few OMR errors that went undetected in the previous OMR stage.

Additional features could further assist score correction. These include the automatic identification of parallel perfect intervals; marking of imitative passages; and, from a more auditory perspective, playback functionality. The latter would especially benefit users who depend more on aural judgment than theoretical analysis.

5.3 Interoperability between MuRET and MP Editor

The integration of MuRET and MP Editor succeeded in automating the transcription pipeline to the fullest extent currently supported by both tools. Future enhancements in MuRET—such as batch processing and improvements to the agnostic‑to‑semantic translation discussed above—will help streamline the workflow even further. Similarly, the improvements discussed for MP Editor will further facilitate more efficient identification and correction of scribal errors.

However, a crucial aspect that remains to be addressed is the level of interoperability between the two tools. After generating a parts‑based MEI file in MuRET, this file must first pass through MP Editor’s input editor before reaching its score editor. During this process, the MEI file is converted into Humdrum **mens—the internal representation used by the input editor. From these **mens encodings, MP Editor reconstructs a new parts‑based MEI file, which is then forwarded to the score editor to produce the final score‑based MEI.

This round‑trip conversion inevitably leads to some loss of information. Features not supported by MP Editor’s implementation of **mens—such as stem direction or fermatas—are discarded.23 While the **mens format itself can represent many of these elements, the subset implemented in MP Editor is limited to the data it can currently input and process (e.g., pitch, note shape, coloration).

An ideal solution would be to allow external MEI files, such as those generated by MuRET, to bypass MP’s input editor entirely and be loaded directly into its score editor. This would preserve the complete structure of the encoding and provide immediate access to MP Editor’s correction and alignment tools, without unnecessary translation steps or data loss. Finally, to ensure consistency and interoperability with mei‑friend, the encoding of IIIF URIs in the MEI files exported by MuRET and MP Editor needs revision.

The MuRET + MP Editor transcription pipeline, excluding the DF, is currently applicable to the principal forms of mensural notation used between the 13th and 15th centuries. Future work on MP Editor’s scoring‑up functionality to support other notation styles would extend the pipeline to work with regional notations (e.g., Italian and English mensural notation).

6 Conclusions

This article introduced the first pipeline for the semi‑automatic transcription of mensural polyphony into fully interpreted, vertically aligned symbolic scores, filling a major gap in existing OMR workflows by integrating symbol recognition, rhythmic interpretation, part alignment, and editorial correction. Although recognition of individual parts may resemble monophonic OMR, aligning them to reconstruct vertical sonorities is a challenge specific to polyphony—one this pipeline successfully addresses.

Built from established tools—MuRET, MP Editor, and humlib’s analyzers—the system reduces manual effort while producing output suitable for both musicological analysis and modern performance. A key contribution lies in reusing and improving existing software while ensuring interoperability, rather than creating a new monolithic tool. This modular strategy encourages extensibility, maintainability, and community collaboration.

The symbolic scores go beyond mere symbol recognition in individual voices. While the retrieved OMR parts already support applications such as content‑based search and concordance comparison, interpreted symbolic scores enable audio playback, automatic conversion to modern notation, and manual or computational contrapuntal analysis. The pipeline thus enhances research and broadens access to the repertoire for performers and audiences.

While this marks a major step forward, future enhancements—such as batch processing and improved agnostic‑to‑semantic conversion in MuRET, playback in MP Editor for counterpoint error detection, and more seamless interoperability between MuRET and MP Editor—will further improve automation, usability, and accuracy.

Demonstrated through applied use, the pipeline shows that this complex transcription task can be largely automated, with output directly applicable to heritage preservation and musicological research. It lays the foundation for scalable, digitally enabled engagement with Renaissance and late Medieval polyphony.

Acknowledgments

Special thanks to David Rizo, Jorge Calvo‑Zaragoza, Juliette Regimbal, Karen Desmond, Craig Sapp, Werner Goebl, and the mei‑friend team for their collaboration; to Ichiro Fujinaga and Julie Cumming for their guidance throughout the project; and to Ellis Reyes for proofreading the complete symbolic corpus prior to its release. Thanks are extended to the Archivo Histórico Arquidiocesano de Guatemala for granting permission to access and digitize the music manuscript.

This paper was written within the framework of the project Echoes from the Past: Unveiling a Lost Soundscape with Digital Analysis (2022.01957.PTDC).

Reproducibility

The dataset can be found at:

Funding Information

This research was supported by a Fonds de recherche du Québec – Société et culture (FRQSC) doctoral grant (2019‑B2Z‑261749).

Competing Interests

Previous PhD supervisors: Ichiro Fujinaga and Julie Cumming.

Colleagues: David Rizo, Jorge Calvo‑Zaragoza, Craig Sapp, and Karen Desmond. They were collaborators on the work detailed in this publication.

Members of the MEI Board and Mensural MEI Interest Group co‑chair: Anna Plaksin, Anna Kijas, Johannes Kepper, Stefan Münnich, Laurent Pugin, David Weigl, Jessica Grimmer, Benjamin W. Bohl, and Maristella Feustle.

Notes

[11] The Verovio Humdrum Viewer (VHV) can be found at https://verovio.humdrum.org, and its documentation can be found at https://doc.verovio.humdrum.org.

[12] For the complete list of filters in VHV, see https://doc.verovio.humdrum.org/filter/.

[15] Since MuRET’s semantic encoding is given in **mens, one can manually add the tokens ‘p’ and ‘i’ to the notes to indicate their ‘perfect’ and ‘imperfect’ values.

[16] All additions to the **mens format were made in consultation with Craig Sapp and David Rizo, as these changes would have implications for both Humdrum and MuRET. The character to encode coloration in **mens was suggested by Rizo.

[17] Work conducted by Craig Sapp.

[21] Check the work on automatic identification of parts by the OLR DeNotEM project https://github.com/Biblissimacluster6/Beyond-DIAMMtoIIIF-DeNotEM.

[22] For future updates on the project and its corpus, consult https://martha-thomae.github.io/projects/guatemala.html.

[23] Also the bounding boxes of individual musical symbols are not transferred to the parts‑based MEI file generated by MP Editor, as this file only retains bounding box information at the staff‑region level. Furthermore, the score‑based MEI files exported by MP Editor do not include any bounding box information, since facsimile data are provided exclusively in the parts‑based MEI output. For the purposes of constructing the final corpus, staff‑level bounding box information was therefore reintroduced.

Additional File

The additional file for this article can be found using the link below:

Supplementary File 1.

MEI File Structure for Encoding a Collection of Parts vs. a Score. It illustrates the structure followed in this project when encoding score‑ and parts‑based Mensural MEI files. DOI: https://doi.org/10.5334/tismir.292.s1.

DOI: https://doi.org/10.5334/tismir.292 | Journal eISSN: 2514-3298
Language: English
Page range: 293 - 308
Submitted on: Jun 25, 2025
Accepted on: Apr 2, 2026
Published on: Jul 15, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Martha E. Thomae, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.