Skip to main content
Have a personal or library account? Click to login
Beyond Ba(t)ch: A Multimodal Meta‑Corpus of Digital Chorale Transcriptions and Related Data Cover

Beyond Ba(t)ch: A Multimodal Meta‑Corpus of Digital Chorale Transcriptions and Related Data

Open Access
|Aug 2026

Full Article

1 Introduction

1.1 Protestant hymn settings

The hymn settings of the Protestant tradition, produced from c.1570 onwards, are well suited for broad digital–analytical examination. These melodies and harmonisations were highly important across a range of musical practices, including public worship and school singing (in German‑speaking regions mostly performed by select ensembles from the local Latin schools). Later, this included congregational singing accompanied by the organ. This was core to a liturgical practice centred on extensive repetition.

This musico‑aesthetic practice was an integral part of the practice of faith, at least in Protestant tradition. The music used in these church services and home devotions exhibited an extraordinarily strong continuity well into the 19th century, remaining much more consistent than other types of multi‑voice music.1

Partly because of that ubiquity, the hymn setting has also been an integral part of traditional four‑voice music and the pedagogy thereof, both in modern music theory and its historical forebears. Generations of musicians received their first instruction in music theory and composition through this repertoire and especially the so‑called Klangfolgetechnik, a technique of controlling chord‑like sonorities which extends beyond a focus on the harmonic triads to include all non‑harmonic motion, such as suspensions. According to this method, different sound combinations (‘progressions’) had to be composed. Repetition and variation were essential didactic features of the method.

Notwithstanding these historical continuities, the chorale tradition also exhibits a degree of complementary range which can make it difficult to grasp at first glance, and concomitantly difficulty to process computationally at scale. There are many aspects to this, including multimodality, which has been a core facet of this genre, even historically.2 The practice of hymn singing has ranged from unison congregational singing to four‑voice school singing (for example, the so‑called Kantionalsatz) and organ‑accompanied congregational singing, and the historical documents exhibit a wider range still (as discussed below). More complex still is the musical framework of home devotions, whose musical forces were always situation‑dependent.

If melodies and harmonisation‑settings can already be regarded as two different ‘modes’ of hymnody, then the settings side of this further divides into at least two further ‘modes’. On the one hand, there are the more fixedly notated sources for a set number of voices (often four), with exact notation (at least of pitch, rhyhm and text) in each. This is often referred to as the ‘Kantionalsatz’, at least for repertoire up to c.1630 and the emergence of the figured bass in this practice.3 On the other hand, there is the more interpretive form with outer voices only (broadly, ‘soprano’ and ‘bass’ in modern parlance), along with figuration. This was clearly intended for the organ accompaniment in congregational singing, a practice that slowly became common after 1600.

To illustrate this range, one goal of the meta‑corpus introduced here is to offer as many versions and variants of different chorale settings as possible and from a relatively broad temporal period.4 In addition to transcriptions of the movements themselves, the meta‑corpus includes related data in the form of texts, basso continuo information, metadata and (where possible) also links to recordings and cross‑referencing of melodies shared across works.

1.2 Motivation for music information retrieval (MIR)

This project concerns itself with the digitisation of historic documents for both preservation and digital interoperability. Chorales have long been a staple of MIR applications. There has been a long history of MIR publications on chorales, though almost entirely limited to J.S. Bach’s settings (e.g. Ebci˙oğlu, 1990; Gabura, 1970; Ju et al., 2020a; Liang et al., 2017; Phon‑Amnuaisuk, 2012; Rohrmeier and Cross, 2008; Winograd, 1968). While chorales provide stylistic range, they do so within a strictly unified remit, for instance, by adopting simpler and more regular rhythms than most musical repertoires (as discussed in Gotham, 2017). This is part of the historic attraction to the Bach chorales as a test case for computational processes as discussed in many MIR projects (e.g. Liang et al., 2017). In addition, the fact that a figured bass is often provided makes chorales a useful test case for simple forms of automatic harmonic analysis.5

Likewise, this repertoire provides unusually clear data for the analysis of phrases, such as their length and contour. Segmentation into musical phrases often has to be operationalised approximately (e.g. by rests) or encoded manually, as in Hauptstimme dataset. Chorale sources, by contrast, typically delimit phrases explicitly in the source using notational elements such as fermatas in the music and punctuation marks in the text (e.g. see Figure 1). These phrases are even of roughly equal length (in terms of textual syllables and musical notes), further enabling simple cross‑comparison among them.

Figure 1

The start of Niels Schiørring’s Choralbook as an example of figured bass realisation.

Chorales also offer clear separation of voices including the ‘primary melody’. In other contexts, this analytical observation has to be manually annotated.6 These data can be used for many tasks, including source separation (on which, see Balke et al., 2025).

Beyond these technical, within‑music details, the data provide wider musico‑cultural insights. These include historical insights to be found in the region‑specific and inter‑personal networks involved in their use for (ecclesiastical) practice and musical pedagogy.

In addition to the above, there are other aspects related to chorales but which are more dependent on the meta‑corpus developed here. First, in focussing on a genre (chorales) beyond the work of a single composer, we create what might be termed a ‘cross‑sectional’ corpus for that genre.7 It stands to reason that these collections provide deeper and more balanced insights into the practice than can be expected of those delimited by a single‑composer. While each constituent corpus of this meta‑corpus is organised in relation to a single‑composer, overall, there is useful balance among them.

Second, unlike most corpora (single‑composer or otherwise), this chorale genre is distinctive in featuring extensive re‑use of the same melodies, with c.100 such cantus firmus parts serving as the basis for chorales. Re‑use of the same melody in different scores throughout a meta‑corpus offers the possibility of tracing historical development and edits over the generations and regions in question. This story is complicated by the fact that these melodies do not always appear under the same title; their development is directly linked to the use of different texts. This phenomenon is examined in more detail in Section 2.4 as part of a wider introduction to the data (Section 2).

2 Chorale Data: Current and Future

In this section, we discuss types of data relevant to chorale corpora. We survey the possibilities in general and what we have delivered in practice (to date, summarised in Table 1). Finally, we provide both ideas and practical guidance on how this collection might be extended in future. For the avoidance of doubt: many but not all elements are shared across the constituent corpora of this meta‑corpus. Not all sources encode the same data, and, indeed, not all elements are appropriate in all cases. This section introduces those issues, mostly in general terms. For corpus‑specific topics, see Sections 4 and 5, and, for the latest file‑specific details, please refer to each corpus’ README file. We begin with the minimum point of complete consistency across all corpus: each chorale comes with a MuseScore file.

Table 1

Summary of data types by corpus (at the time of going to press)

CorpusSymbolic ScoreMetadataText (‘Lyric’)AudioFigured BassCalendar
Goudimel (Section 4.1)41YESYES.N/AImplied
Bach (Section 4.2)389YESYESYESN/AYES
Rein (Section 5.1)201YES..YES.
Zinck (Section 5.2)162YES..YES.
Kittel 24 (Section 5.3.1)24×8+YES..YES.
Kittel Vierstimmige (Section 5.3.2)162YES..YES.
Schiørring (Section 5.4)102×3 versionsYES..YES.
Apel (Section 5.5)177YES..YES.

2.1 Digital scores

Symbolic representations of primary sources (manuscripts, prints) are central to this meta‑corpus. We represent each chorale as an individual file in the XML format of the Free and Open Source score editor ‘MuseScore Studio’.8 This enables users to consult the music notation directly, without complex computational set‑up or other overheads. From this version of record, we export all score elements and their properties as tabular files, which can be processed with any spreadsheet or statistical software tools. In this way, digital representations of the primary sources at hand serve as anchor points for the alignment of other dataset components such as audio files. For increased interoperability of our chorale dataset, we provide the digital scores in two additional formats, .mei and .musicxml.9

2.2 Metadata

In order for scores to be used as research data, a set of clean and consistent metadata is essential. Our digital scores make use of MuseScore’s feature for storing metadata as key‑value pairs directly in the file. We use the pre‑defined keys (such as ‘composer’, ‘workTitle’, or ‘workNumber’) to identify the chorale that each file represents. In addition, we have defined custom keys to embed the scores in a Linked‑Open‑Data paradigm, by including (where possible) work identifiers from the MusicBrainz project.10 For increased accessibility of the metadata, we also provide one TSV file per subcorpus with one row per chorale and one column per key, as well as one global TSV file containing metadata for the entire meta‑corpus dataset. These tabular files contain the respective file names and are therefore instrumental for accessing the data programmatically. In addition, they allow for batch‑editing metadata of many scores at once by modifying the relevant tables and updating the MuseScore files using Hentschel and Rohrmeier (2023)’s score parser ms3.

2.3 Aligned audio recordings

Although it would be desirable to provide free, public‑domain audio recordings alongside all the chorales, such recordings either do not exist or are not currently available under suitable licence for almost any of the items in this collection. Instead, we provide unique identifiers to commercial recordings wherever they exist and where such a matching could be undertaken automatically. Specifically, we used the MusicBrainz database to match up 304 chorales by J.S. Bach with recordings by the Chamber Choir of Europe and the Freiburger Barockorchester, recorded in 1999 (Chamber Choir of Europe et al., 2014). For these 304 recordings, we provide alignment information for each note of the four choir parts in milliseconds, as computed by the Dynamic Time Warping function in the Sync Toolbox (Müller et al., 2021). Through this direct correspondence of musical time (each note’s position in the score) to absolute time (the corresponding position in the recording), it is possible to align other score elements (such as lyrics) with the recordings as well. Although these alignments support many research possibilities in their own right, researchers wanting to make full use of the recordings will need to acquire those recordings directly.

Apart from augmenting the number of linked commercial recordings and the provision of public‑domain recordings, future work could include further modalities such as videos of chorale performances, likewise aligned in this way. This could include both ‘stand‑alone’ chorale recordings and those appearing as a constituent part of a longer work e.g. a Bach passion, along with timestamps for where a specific chorale starts and ends.11

2.4 Text, verses, meter

Texts are of critical importance in church music. Here, we include chorale texts where they are present in the source: for Goudimel and Bach. As discussed above, it is a typical feature of church music, however, that a chorale melody can have more than one matching text. A melody can be assigned to a different text (and title) in different chorale books, which can result in differences in the music – for example, in the number of notes and even fermatas. There are around 100 chorale melodies that appear in chorale books right across the centuries and across similarly wide geographic spans. Finding those melodies in different chorale books and then comparing them (in terms of music and text) can offer a deeper insight into compositional conventions in church music in different times and places. The meta‑corpus contains a general overview of all constituent datasets (including metadata) in which we link scores known to have matching melodies to each other by assigning them a common Set‑ID.12 This is accompanied by a list of these Set‑IDs with their corresponding various titles and allocations in the corpus. We provide some computational scripts for attempting the automation of this as part of a set resources of shared utility in a dedicated analysis repo on the Chorale Corpus.13 Naturally, this involves distance measures that have some tolerance for extensions (such as a coda appended to one instance and not another) and for discrepancies in the use of passing notes, for instance (by comparing reductions rather than the original). Such routines never perfectly avoid all false positives and negatives, at least not across different corpora. For examples of operational approaches taken, see the discussion of clausulae pairs in Remeš et al. (2025).

For texts and tunes broadly outside of this ecosystem, we provide complementary data. For the Goudimel corpus, for example, we provide the first verse of the text for all transcribed files in stand‑alone text files (.txt). This is complemented by wider resources for navigating psalm texts and numberings. We provide (for what appears to be the first time) side‑by‑side numbering of the two main numbering systems (Greek and Hebrew) and code for mapping between them (as directly as possible), ‑along with a complete English‑language version of the text (King James Version ‘KJV’) and titles for each in French (Goudimel) and Latin. This has proven useful for spot‑checks and has clear utility for wider data‑driven work on the psalms beyond this collection, and beyond music.

2.5 Figured bass (a.k.a. basso continuo)

‘Figured bass’ is a form of musical notation whereby the bass line is furnished with symbols that indicate pitches to be played in upper voices. The symbol syntax is based on numbers that relate to generic intervals (e.g. ‘3’ for a diatonic third above the bass note), which can be modified to indicate accidentals (specific intervals) and combined freely. While compound intervals were sometimes notated (e.g., ‘10’) the latter, more standardised practice usually saw octave equivalence assumed (so ‘10’ becomes ‘3’).

The figured bass line may stand alone, or may be combined with one or more fully written‑out upper line (which should then correspond to the implication of those figures). As with much semi‑analytical notation, figures under‑determine realisations. The (visual) Figure 1 provides a highly useful and relevant example from the meta‑corpus in which the composer‑pedagogue sets out ‘both’ of these realisation types. The lower grand staff features only the outer two voices (‘soprano’, ‘bass’), the bass of which is figured. The upper grand staff consists of a four‑voice realisation with these two (‘soprano’, ‘bass’) as well as two inner voices (‘alto’, ‘tenor’), which realise the figures.

Beyond the functional role of providing generic harmonic information for the accompanist, figures also offer an authentic insight into the contemporaneous ‘Klangfolgetechnik’ (on this term, see Section 1.1). The comparison of different figured bass conventions across time can reveal significant changes in the use of this technique. In this meta‑corpus, we aim to provide machine‑readable figured bass wherever possible to support studies of this kind, though this poses significant challenges, as discussed in Section 3.1.

2.6 Function and position in liturgical calendars

Many individual chorale melodies have a close connection to a specific occasion or wider season of the liturgical year. Moreover, many (perhaps most) of these connections are as old as the melodies themselves. After the Reformation, such traditions were consciously continued and expanded. The church year (which, in the Protestant church, begins with ‘Advent’) therefore also became a central organisational structure for the Reformed church’s renewed collection of hymns, which in turn was reflected in the arrangement of the hymns in the surviving source materials. For example, many older collections begin with the hymn Nun komm der Heiden Heiland, which marks the start of Advent. Later collections see these traditions replaced by alternative organisation systems, such as alphabetical ordering by the first line of the text. In order to understand this historical development of the correspondence between melodies and the seasons of the liturgical year, it is therefore necessary to consider the position of the melodies in the church calendar, their assignment to specific texts and even the semantics of the texts themselves.

The liturgical year is not the only relevant temporal cycle here. Our collection includes Goudimel’s setting of the psalm texts (Section 4.1). The liturgical use of psalm texts is more typically attached to a position defined by the time of day and/or month rather than the year. In the Goudimel source, the introduction includes a double index to this effect: the first is alphabetical, and the second allows the reader to look up the psalms according to ‘the order in which they are sung in the Church of Geneva’.14

2.7 Dataset‑curation infrastructure

We describe the present contribution as a ‘meta‑corpus’ because it constitutes both a number of individual corpora (which can readily stand alone) along with a degree of aggregation into something consistent enough for use as a many‑corpus whole. Each sub‑corpus constitutes a single ‘work’ (typically chorale book) by a single composer and is both stored as a separate GitHub repo (within a shared ‘org’ dedicated to this meta‑corpus) and indexed on the Zenodo platform with its own DOI. The meta‑corpus is aggregated into another repo on the same org.15 We employ GitHub’s continuous‑integration features to automatically release a new version and mint a version‑specific DOI each time a sub‑branch is merged into a repository’s main branch following a merge request. Through this meta‑corpus structure and Zenodo linking, we provide the means to download either individual corpora or the entire meta‑corpus at once, both as a repository containing git sub‑modules and as a frictionless data package consisting of four TSV files concatenating the data from all sub‑corpora (see Figure 2). This infrastructure mirrors that of the Distant Listening Corpus, described in detail in Hentschel et al. (2025).

Figure 2

Original encoding formats of the eight sub‑corpora and their conversion paths to the authoritative MuseScore files (.mscx) that each of them is provided in. .json: Tiny notation; .mxl: compressed .musicxml; .mscx: Uncompressed MuseScore 3.6.2 format; .xlsx: Excel format; .marcxml: MARC 21 metadata format; .mei: MEI Basic 5.0, converted with MuseScore 4.4.

3 Shared Encoding Issues

3.1 Figured bass

Figured bass is challenging on many fronts. First, while the practice of figuring was common throughout the time period represented by this meta‑corpus (and was used well beyond chorales), it was neither consistently used in all the sources represented here (see Table 1) nor used consistently across those cases where figures are present. There are divergent regional and historical practices in both the notation and realisation of figured bass.16

Among differences in the notational practice, while a relatively consistent practice of sorting the numbers in numerical order emerged (with ‘top‑to‑bottom’ corresponding to ‘largest‑to‑smallest’), this was by no means a consensus even in 1800. Figures were often arranged according to the voice‑leading of upper parts (which may also be written out in the notation). It is precisely such figures, however, that often seem to echo a notational practice (common in chorale books of the 18th and early 19th centuries) with only melody and figured bass. The transmission of figures is therefore problematic: here, as so often, different use cases call for divergent solutions. If the aim is to support practical readability for conventionally trained modern performers and/or enable direct comparison for computational–quantitative evaluation, we might be inclined towards an overarching standard. By contrast, if the goal is an investigation into the historically prevailing conventions in the practice of figuring itself, the original layout must be preserved. Under certain circumstances, the original notation can elucidate conclusions about the origin and context of the material. Original basso continuo parts are also likely to be informative with regard to a number of technical compositional conventions that are otherwise no longer traceable.

Turning to a concrete example, even given only the figured bass line and no other information, it is clear to practitioners that, in Figure 1, every bass note should be harmonised, including those without a figure (which are taken to be ‘53’ where not specified)17 but excluding the one instance of quarter note motion where the C is understood to be a non‑harmonic passing note. Both principles are almost but not quite generalisable as rules, and as the rhythmic density of the music increases, so too does the potential ambiguity.

Related, the duration (or equivalently, ‘end’) of a figure is rarely specified, and even the starting position of a notated figure can be ambiguous. While many figures align clearly with a bass note (and are understood to start at the same time), there are many cases of upper voice movement against a static bass, for instance, where the horizontal positioning of the figures during the given bass note does not unequivocally identify a metrical placement. Solutions depend on the performer’s instinct and experience: even if the intended voice‑leading detail can, in principle, be expressed with two figures, the possibilities of fixing such a movement with rhythmic precision are clearly limited. An 8–7 motion over the static bass of a cadential dominant chord is especially common, and there has never been a clear standard for such situations, even though attempts have been made in this direction many times.18

The chorale repertoire is more amenable in this respect than other, more rhythmically complex genres: due to the syllabic text setting and relatively simple rhythms, chorales have a high density of figures (relative to the number of notes) and a higher degree of alignment clarity in general. Suffice to say, digital analysis of the basso continuo is still in its infancy, and chorales offer a good entry point.19 As so often, computational encoding and analysis shine a suitably uncompromising light of these issues, and are perhaps best tested in the relatively simple context of chorales before expanding.

Given the lack of a clear and consistent reading of historical basso continuo, it is perhaps unsurprising that notation programs and analysis tools are similarly inconsistent and un‑interoperable. While formats like .mei and .mscx can encode figured bass, they do so differently. Friction and failure are to be expected when converting, particularly for figures without a direct reference note.

This leads to ad hoc workarounds. One way to achieve a correct conversion of the figured bass from MuseScore to .musicxml or .mei is to add the figures to an additional (‘fake’) voice with notes at all required positions. Examples for this method can be found in the Chorale Corpus GitHub but are intended as a proof‑of‑concept: while this is somewhat practical, it is clearly far from ideal.

3.2 Text string encoding of early music

As discussed, we provide MuseScore files as the file format point of consistency across this meta‑corpus. The advantages are clear; however, there are also some disadvantages, including the over‑specificity of a modern digital score in relation to the source material. For some of the datasets, we provide a demonstration of an alternative or ‘complementary’ encoding: a text‑only representation of the source from which the MuseScore and other versions can be generated.

Text string encoding of early music confers various advantages. First, it is an extremely light‑weight way of representing scores. The string itself is also maximally interoperable: everyone can open it with no non‑standard software, though the conversion from string to score rendering requires more specialist handling.20

Perhaps the main advantage of string encoding for early music is its edition‑ and version‑neutrality: encoding musical scores in most standard notation editors requires commitment to time signatures, key signatures, voice ordering, and more. Encoding in text strings, by contrast, allows a representation that is more neutral and which arguably reflects the original source material more closely.

Text strings and the separation of values also enable more human‑readable version control. For example, changes to a single note can be deciphered more clearly in the context of surrounding notes than when buried in layers of XML. This also allows easy and immediate rendering of distinct ‘ancient’ and ‘modern’ versions, as discussed in Section 4.1.21

While the text strings are a promising way of encoding scores, a possible detraction is that they are less intuitive as an input method. This is particularly relevant where crowd‑sourcing is involved: ‘the crowd’ is typically well versed in a notation software editor like MuseScore and none of these alternatives.22

This leads to another possibility for using string‑based encodings, as a ‘downstream’ derivative from the (notation) version of record. That is, transcription takes place in any convenient form, and perhaps even from more modern sources, and then the corpus hosting effort reverse‑engineers the process, creating the string (e.g. .txt or .json) from the original transcription (e.g. mscz). The original format can then continue to serve as the version of record, while the ‘off‑score’ version provides clearer version control and rendering options.

In this meta‑corpus, MuseScore is the central (shared) format; in most cases the files originated there and all other formats are extracted from it. In just one case (Goudimel, Section 4.1), we work in the other direction, building those MuseScore files from strings. And, in some others, we provide the string representation as a derivative. Overall, we see significant prospects for leveraging crowd‑sourcing encoding flexibly; this is particularly relevant for early music that is not weighed down with extensive annotations of dynamics, articulations and more.

3.3 Dataset creation process

As with the data, so too with the creation process: the constituent datasets have been created by different groups with divergent perspectives and interests; these datasets are brought together by an inter‑group (and international) collaboration. The benefits of coordination and scale will be clear to the MIR readership, and the challenges that come with this coordination are probably just as clear.

Since each data creation/curation effort was well under way before this collaboration began, the processes for each were quite divergent (see Section 5). To consolidate the data, we therefore agreed on a common target format (uncompressed MuseScore format, .mscx), which we reached in different ways. This can be retraced through the chart in Figure 2.23 The chart shows for each sub‑corpus the first digital format that it was encoded in and the corresponding conversion steps to the target format .mscx (many of the sub‑corpora originate in that format and did not require conversion). In addition, an .mei file is generated from each .mscx file.24

The metadata is also extracted into an additional separate file. For coordinating consistent display, metadata are extracted into a tabular file (.tsv) in each sub‑corpus. Updated metadata can also be written from the .tsv file back into the .mscx file. For some corpora, the metadata has been included in the RISM‑OPAC, and these checked entries were exported to the cataloging format .marcxml (sic) to update the original metadata. All told, despite the divergent original creation processes, all transcriptions are therefore published in the same open format, with their metadata listed in a uniform tabular file.

4 New Corpora

This and the following section (Section 5) outline the constituent corpora of our meta‑corpus, along with issues specific to each. In this section, we introduce two ‘new’ corpora in the sense of not having been previously published and discussed in any existing dataset paper: Goudimel (Section 4.1), and Bach (Section 4.2). Being new, these warrant a longer explanation; Goudimel is also the most different from the others.

4.1 Goudimel, Claude (c.1514–1572): The Geneva Psalter, 1564

The Goudimel corpus is an outlier in many respects, complementing the other corpora in terms of text (setting the psalms), language (in French rather than German), and time period (being by far the oldest).

Our source for this transcription is the second edition of 1564: the earliest source on IMSLP, hosted there in two files. The two files are for the first half (1–68)25 and the second half (69–150).26 IMSLP also provides a comparison of sources using the Genevan Psalter melodies.27 This original source is set out in the ‘Choirbook format’ that was common at the time, allowing all four voices to read from the same double page spread. This layout is neither like the modern score (all voices vertically aligned) nor modern parts (all voices on separate documents). Perhaps the closest modern equivalent is the four‑hands piano format where two players (primo and secondo) play on the same piano and read from the same double page.

Partly because of this format, the Goudimel corpus lends itself to the string‑encoding of music introduced above in Section 3.2. As mentioned there, this is the only case in which the string encoding is the version of record from which all secondary sources automatically derive (most uses extract the string representation as a derivative).

In this corpus, we experiment with a bespoke string‑encoding solution. The goudimel.json file contains all included score and metadata (title, psalm number, ...) in one place.

All 150 items include minimally:

  • original bass: a pitch name with octave string (e.g., “C4” – not to be confused with the clefs) denoting the first bass note.

  • original key: The original key signature expressed as a bool: there is either no key signature (False to indicate zero sharps/flats) or a key signature of a single flat (represented by True).

  • original clefs: a list of four entries, specifying the sign and line of the clefs in the original parts (from superius to bassus), e.g., ‘G2’, ‘C3’, ‘C3’, ‘F4’.

For those that have been transcribed fully, the .json file includes the string encoding of the music and (separately) key:value pairs for relevant parts of the musical source that are not readily included in music21’s tinyNotation:28

  • modern transposition: the number of semitones and direction (+/‑) for transposing from the original to the modern key choice.

  • modern sharps: the key signature of a modern version expressed as a number of sharps. (Note that this cannot be deduced from the transposition as we are dealing with modal sources.)

From the .json data, we provide code for rendering scores in both ‘original’ and ‘modern’ versions. ‘Original’ here means that it follows the source in terms of the:

  • ‘open score’ layout, with one voice part per stave, (though the original is in the “Choirbook format” as discussed above)

  • part names (‘superius’, ‘contra’, ‘tenor’, ‘bassus’)

  • clefs: including all original C‑clefs like ‘soprano’, ‘alto’, ‘tenor’. The usage varies across the source. Each part sees a limited range of clefs across the corpus: each voice range makes use of two options across the corpus except the bass which uses three. There is also a very uneven use of the resulting combinations. For instance, the combination C1–C3–C4–F4 appears 88 times, while the seemingly small variant of C1–C3–C4–F3 appears only once.29

  • note values, moving mostly in whole notes.

  • part distribution with the cantus firmus melody placed in the tenor (middle of the texture) as per Goudimel’s original harmonisation.

  • transposition level.

The original is almost entirely expressed in the .json file, with the following notes, that could be viewed as workarounds, or design choices:

  • Final notes are usually double the length: the longa note value (as opposed to breve) is unsupported by tiny notation.

  • The clefs are in the .json file, but separate from the tiny notation, both because clefs are not supported as standard and because storing the clefs separately helps with certain tasks like the comparison discussed above.

  • time signature symbol: tiny notation supports the concept of 2/2, but the symbol for ‘cut time’ is added later.

The ‘modern’ scores:

  • retain open score layout (though with a script provided to adapt to short score),

  • adopt modern part names (SATB) and modern clefs (treble, treble, treble8, and bass),

  • move the cantus firmus from the tenor to the top‑most voice,

  • halve the note values, and

  • transpose by the amount and direction indicated.

Modern editions vary. This is part of the motivation for the flexible string encoding: the choices here are easy to alter. For example, the transposition is a single integer. This can be changed directly in the .json file and the score re‑rendered programmatically. That said, there are common practices such as the mapping to SATB and the assignment of modern clefs. Where the choices are less clearly‑defined, we take options that serve to align this corpus of scores with the corresponding harmonic analyses by Tymoczko et al. that are hosted on ‘When in Rome’.30 This alignment is part of the motivation for including Goudimel here, and our Chorale Corpus deliberately covers all items included there.

4.2 Johann Sebastian Bach (1685–1750): 389 Choralgesänge

J.S. Bach’s four‑voice chorale settings are the most well‑known among those represented here. They continue to play an important and prominent role in the modern‑ day, in the concert hall, in educational settings and in computational music research. The relative homogeneity of chorales in general and the notoriety of Bach’s chorale settings specifically explain in part why they constitute one of the oldest and most widely used of computational corpora, and why they have been circulating in many different formats for decades. Compared to the pre‑existent Bach Chorales datasets – which typically come without lyrics and comprise 370 chorales using Riemenschneider’s catalog numbers – our collection contains all lyrics and reflects the more recent Breitkopf 1912 edition from the Public Domain, entitled 389 Choralgesänge. Thanks to its handy format, its low price and the alphabetical ordering, this resource is popular with musicians and music scholars alike and has come to be known as the ‘purple bible’.

Today, many regard Bach’s chorales as classic representatives of their genre. This assessment is at odds with their relatively low significance for 18th century congregational singing. At that time, Bach’s movements were considered esoteric, and their significance for the historical liturgy, worship and service was certainly very limited.31 The undeniably high quality from a compositional and musical point of view has both ensured their longevity and tended to obscure a balanced view of wider hymnal practice from the time.

5 Integrated Corpora

In this section (Section 5), we (re‑)introduce more briefly than in Section 4 those corpora which have been previously described and published in some form – Rein (Section 5.1), Zinck (Section 5.2), Kittel (Section 5.3), Schiørring (Section 5.4), and Apel (Section 5.5) – with the main previous works being Gerhardt and Kirsch (2024) and Remeš et al. (2025).

5.1 Rein, Johann Balthasar (1714–1794): Vierstimmig Choralbuch

The Vierstimmig Choralbuch, worinnen alle Melodien des Schleswig‑Holsteinschen Gesangbuchs enthalten sind contains 201 outer‑voice chorale settings with figured bass (202, if second versions are counted individually). The texts corresponding to the titles are not part of the original chorale book, but we strive to add them to the corpus. Rein’s chorale book also contains an overview of which texts can be sung to the melodies, since more than one text can fit the same melody. This information will be added to the metadata of each score, since it can be an important aspect when analyzing a chorale. In order to handle the problem that conventions for dealing with the figured bass can interfere with a computational analysis – e.g. if the resolution of a fifth has not been explicitly notated – we added a second version of each chorale, in which these possible sources of error have been adjusted.32

5.2 Zinck, Bendix Friedrich (1715–1799): Vollständige Sammlung der Melodien

In the Vollständige Sammlung der Melodien zu den Gesängen des neuen allgemeinen Schleswig‑Holsteinischen Gesangbuchs, Zinck presents 138 different outer‑voice chorale settings with figured bass (162, if second versions are counted individually). As for the Rein, we plan to link texts corresponding to the titles directly to the score and the information about possible alternative texts from the register to the metadata. We have also added a second version of each chorale for this book that contains changes needed only for an error‑free digital analysis (as discussed in Section 5.1).

5.3 Kittel, Johann Christian (1732–1809)

5.3.1 24 Choräle

Johann Christian Kittel’s 24 Choräle mit acht verschiedenen Bässen über eine Melodie is a set of 24 chorale tunes, each of which is harmonized by eight to nine distinct bass lines that broadly increase in both complexity and rhythmic activity from the top to the bottom of the page. This graduated difficulty serves a pedagogical function, as described in Remeš et al. (2025).

Of the 24 melodies, 22 have eight basses and two have nine basses, amounting to 194 melody–bass pairs overall. As many chorales survive only in this melody–bass format and as the re‑use and re‑harmonization of cantus firmus chorale tunes is a core part of the Protestant chorale style, this amounts to a corpus not dissimilar in scale to the Bach, for instance.

As with the Goudimel, we provide this corpus in several configurations. While Goudimel is presented in ‘ancient’ and ‘modern’ versions, the main distinction here is between 1) separate scores for each melody–bass pair, and 2) each melody combined and aligned with all eight to nine bass lines to facilitate visual comparison (not to be played back at once!). We also include the text‑only version as described for Goudimel above (though in this case, that format is converted‑to).

5.3.2 Vierstimmige Choräle mit Vorspielen

Kittel’s Vierstimmige Choräle mit Vorspielen contains 155 chorale settings (162, if second versions are counted individually). All are in four voices, complete with figured bass, and preceded by an instrumental prelude. So far, only the chorales themselves have been transcribed as outer‑voice settings with figured bass. We plan to digitize the middle voices and the corresponding texts (according to the chorale’s title) in the longer term. As for the Rein, we plan to link information about possible alternative texts from the register to the metadata. We have also added a second version of each chorale for this book that contains changes needed only for an error‑free digital analysis (as discussed in Section 5.1).

5.4 Schiørring, Niels (1743–1798): Choralbook

Danish composer and theorist Niels Schiørring’s Choralbook sets out each of 102 chorales with three forms of harmonization/realisation: 1) a four‑voice choral setting; 2) a four‑voice realisation intended for the keyboard (using the grand stave, two‑hands); and 3) a two‑voice keyboard setting with figured bass to indicate the inner parts. Figure 1 provides an example of items 2 and 3.

As with the case of Kittel’s multiple bass chorales (Section 5.3.1), we provide a text‑only version and also combine/render the various components as separate scores for each of the versions described above, as well with the vertical combination to aid visual comparison.

5.5 Apel, Georg Christian (1755–1841): Vollständiges Choralbuch

Apel’s chorale book Vollständiges Choralbuch zum Schleswig‑Holstein’schen Gesangbuche consists of 177 chorale settings in four voices. We transcribed these settings true to the original and added a second version in which we added a figured bass. This way, we are able to compare the settings to other chorales in our corpus that contain a figured bass because we can still use the same tools. As for the Rein, we plan to link texts corresponding to the titles directly to the score and the information about possible alternative texts from the register to the metadata.

6 Outlook

This outlook section concentrates on a number of possible extensions to this data‑gathering project. We focus primarily on prospects for expanding the remit to related genres (Section 6.2) and to later historical periods (Section 6.3). Before that, we begin with a few comments on how the existing data could be enriched.

6.1 Data enrichment

All aspects of this dataset could be developed both in terms of expanding the provision of data and of retrieving information from that data. At one extreme, a much larger project could obviously extend the multimedia and/or multimodality of the corpus.33 It would be valuable to include new audio and video performances, perhaps with multiple cameras and microphones (see discussion in Section 2.3). Just as obviously, this is a major undertaking well beyond the present scope.

At the other extreme of extracting information from one aspect of the available data, one could derive meaningful insights from the text alone. For instance, given a suitable encoding it is possible to extract the numbers of various elements such as syllables, lines, and verses. Formal encoding options include the TEI standard; simple, and human readable alternatives include encoding text with divided syllables (e.g. with hyphens), one line per phrase, and a line break per verse as demonstrated in the Goudimel corpus. From this, one can explore the textual meter (e.g., 8.7.8.7) both alone and in relation to other considerations such as the prospects for text–tune pairings. Naturally the information in text extends far beyond basic counting of elements. Topic modelling could explore the subject matter and emotional tone of these musico‑textual works, perhaps again with a view to comparison, e.g. for liturgical season. Again, this is dependent on the ability of such tools to parse languages beyond English and in historical styles before the modern era.

Between these two extremes are any number of other project possibilities. For example, one could explore the significance of this ‘everyday’ repertoire for sacred art music in its various forms. It would be interesting to explore historical changes in religious–functional music as it moved towards new roles in the concert hall and thus ‘autonomous’ musical art forms. The next section explores one aspect of this.

6.2 Expansion on the genre level in the 18th century

6.2.1 Oratorio church music

In an oratorio, chorale and hymn settings have different and changing roles with regards to topos, quotation, dramaturgy, and liturgical reference. Consider, for example, the so‑called Passion Cantata of Carl Philipp Emanuel Bach, which adapts a St Matthew Passion originally composed for the main church service. The liturgical version contains pieces by C.P.E. Bach’s father (J.S. Bach), as well as by Georg Philipp Telemann and Georg Anton Benda. For the concert version as a Passion oratorio, C.P.E. Bach removed precisely those movements that were most directly attached to liturgical use. This does not mean the removal of all hymnic items. Two examples from this piece may serve as representative of the field of chorale use in oratorio music (which has seen little systematic investigation to date).

First, in the great chorus Lasset uns aufsehen auf Jesum Christum, the vocal part enters the figural orchestral part in unison with long notes, i.e. cantus firmus‑like as shown in Figure 3. Here, a special melodic–harmonic twist draws attention: the repeated notes and the following modally striking melodic turn A‑C (‘look up’) in the D major context are reminiscent of the well‑known chorale Es ist das Heil uns kommen her.34

Figure 3

Reduction of Lasset uns aufsehen auf Jesum Christum from C.P.E. Bach’s Passions‑Cantate (Hamburg, 1789). The reduction and transcription (by last author, MG) show the choir’s S,A parts (T,B are 8vb) and the orchestra’s continuo part (which is doubled in orchestral unison including at the octave). Non‑doubling orchestral parts enter with the voices (not shown).

It seems clear that there is a contextual reference here, but the choral movement or melody contains no other similarities with the chorale in question. Could Bach have used such a specific turn of phrase here and implemented its context and semantics in his composition? Are there other chorales with similar turns of phrase? How often might such possible references occur in phrases beyond the incipits?

Second, at the central point of the Passion narrative, before Jesus’ words ‘Es ist vollbracht’, C.P.E. Bach sets the only chorale that is also labelled as such: Heiliger Schöpfer, Gott. In fact, however, this is only the second part of the hymn Mitten wir im Leben sind, which was composed by Luther as Heiliger Herre Gott and was common in the Hamburg hymnal, reworked by Balthasar Münter as Heiliger, Schöpfer, Gott (published 1772) and adopted by Bach as Heiliger Schöpfer, Gott in a slightly modified meaning without the first comma. As the first part of the chorale ‘Mitten wir im Leben sind, mit dem Tod umfangen’ has already been ‘dealt with’, as it were, by the Passion action, the choir comes forward with this invocation, which quotes the choral tone most clearly. In this form, despite the simplicity of the movement and the instrumentation and despite the familiarity of the melody, the ritual fulfilment is prevented and the aesthetic and functional roles of the melody are fundamentally changed.

In order to find such references and quotations in sacred choral music, a search of text and melody incipits is not sufficient. Rather, it is necessary to search for text and melody excerpts and variants across the entire repertoire. Using computational tools for this tasks comes with some requirements that differ from those of common incipit searches. First, we need searchable digital transcriptions of as many chorale books as possible – a goal to which this cooperatively compiled corpus is intended to contribute. Second, the search method not only needs to include the whole melodies and settings but it should be adapted to a particular aspect of church music. As discussed, chorale melodies have often been aligned with more than one text and therefore different numbers of syllables, which can lead to further differences such as the interpolation of additional notes. Naturally, this requires subtler search functionality than sheer n‑grams, for example. A separate repository of the Chorale Corpus reported here includes some methods for this,35 including the MusiCAU’s previously reported ‘Phrase Detector’ as well as comparable methods used and reported in Remeš et al. (2025).

It seems likely that this kind of reminiscence carries semantically connotative characteristics. Within 18th century transfers of such musical material, the ‘choral‑topic’ doesn’t seem to be as tightly bound to the compositional technique or characteristic motives. As Janice Dickensheets observes of Classical–Romantic music, these features are only later distilled into the ‘Topical Vocabulary of the Nineteenth Century’ (Dickensheets, 2012).

These aspects might form the focus of future work on the basis of a broader choral dataset, with the aim of gaining more insight into specifics of ‘the chorale‑topos’ within 18th century musical language, and perhaps the projection thereof onto a wider range of genres beyond 1800. Relevant here is Eric Chafe’s highly illustrative study of how J.S. Bach works with decidedly ‘modal anomalies’ in the chorale settings of his cantatas (Chafe, 2000). Future work could potentially engage Chafe’s hypotheses and perhaps identify further, previously undiscovered cases.

6.2.2 Song culture: ‘choral’ songs

The sacred solo song is another contemporaneous genre into which the chorale crosses over in various forms and thus develops a life of its own. Once again, this is a repertoire which has not yet been systematically investigated and which will be more practical to analyse with appropriate preparatory work on the ‘reference genre’ (i.e. the chorale, as here). Most of the pieces in the various printed collections of songs and odes (sacred and otherwise) are ‘empfindsam’ (sensitive) art songs. Examples include Münter’s Erste Sammlung Geistlicher Lieder (Leipzig, 1773), Carl Philipp Emanuel Bach’s Cramer’s übersetze Psalmen mit Melodien (Leipzig, 1774), and Friedrich Ludwig Aemilius Kunzen’s Oden und Lieder (Leipzig, 1784).

These songs are small character pieces and have very different affective forms and gestures. They usually consist of a piano setting in which the upper piano part doubles the vocal melody (as is common in many art song genres). These pieces are often instrumentally conceived piano pieces with a vocal part that is quasi ad libitum. Among them, chorales or chorale‑like songs or individual verses make up a relatively small proportion: in Münter’s collection of songs by various composers (including older ones such as those by Johann Scheibe), 22 out of 50 have a ‘choral’ reference. In C.P.E. Bach’s collection, it is 9 out of 42 and, in Kunzen’s, only 5 from 91. In context, their chordal–homorhythmic setting in large note values usually means they stand apart from other songs of the collection which are typically more richly ornamented. Then again, both share a textural frame of ‘soloistic upper voice setting with accompaniment’.

There are also complex combinations and differentiations of styles and aspects of chorale settings within the songs. For instance in Münter’s collection, song 16 (Lied XVI: Verkündigt alle seinen Tod!) uses two different chorale setting styles, as Johann Adam Hiller combined a first part of each stanza in a ornamental four part setting for a ‘Chor’ and a second rather plain part for a ‘Gemeine’ [sic], that is to say, for a ‘community’ to sing (whether that community is real or figurative). C.P.E. Bach’s Über die Finisternis kurz vor dem Tode Jesu (Sturms Geistliche Gesänge, S. 29) includes the melody of the chorale Christe Du Lamm Gottes but without using the typical homorhythmic ‘chorale style’ setting.

Questions that need to be addressed for further investigations into this ‘chorale‑like’ song repertoire include: Which references do specific examples really have to well‑known hymns? Which aspects of the chorale are referenced here as distinct from more fundamental compositional techniques (e.g. certain clausulae)? Which harmonic turns and phrases occur frequently, and thereby form the characteristic ‘choral‑tone’? Which well‑known texts and meanings might be referenced? A broad access to the reference repertoire of the chorale‑settings is essential for those further investigations.

6.3 Historical expansion into the 19th century (and beyond)

In the above, we have some repertoire contexts in which the chorale seems to expand beyond the liturgical–ritual context and become a topos in itself. This is especially (but not only) relevant to the ‘new’ concertante genres of vocal music. At the same time, a chorale can be the bearer of an important aesthetic moment – namely, a ‘sacred’ devotion, the centre of which, depending on the context, can continue to be a religious one, but which also becomes specific to the ‘art–religious’ attitude of reception in the concert hall. Thus, the chorale and the hymn setting can be regarded as a quotation, an archaic ‘remnant’, a character, an allusion, a topos and a medium of special gravitas.

Here, the texts add a separate level of meaning, whose change in the 18th century against the background of different currents of the Enlightenment though, poetics and theology, should be considered alongside regional and temporal variants. Texts can also semantically adhere so closely to a melody after centuries as to contain ‘concealed’ meanings. As such, they should not be thought of exclusively as musical artefacts in instrumental music either but also in relation to the most relevant texts in each case. (see the discussion of text valence in Section 6.1). Thus, chorales continue to have an effect not only musically but also in terms of content and not only in sacred vocal music but also in secular vocal music and the ‘absolute’ instrumental music of the 18th and 19th centuries and beyond. The integrative power of music, which is now emancipating itself within the arts, becomes particularly clear in the movement of music from previously being bound to rituals (and thus essentially functional) to something ‘freer’.

While our collection reported here extends well beyond Bach, we are still largely limited to the 18th century. And, while chorale collections become rarer (or at least change in nature) as the centuries advance, ‘the chorale’ or ‘hymnal’ style is highly recognisable and present in many later works, with or without direct quotation of a specific chorale tune. In an ‘issue’ at our central data repository,36 we provide a list of some cases. These include both individual works and genre examples of the 19th and 20th centuries. The list should serve to keep a record of ideas and perhaps to inspire extensions.

Future work could attempt to navigate this ‘afterlife’ of the chorale, exploring, for instance, the recurring, thematic uses of chorales (e.g. as a topos referring to death and memorial). In order to better understand these phenomena and understand their transformative function, the links between everyday practice and art music production can be targetted directly. For this, we would need a repository that is as extensive and accessible as possible as well as search options that are as specific and differentiated as possible. Relevant reference values that are already recognisable include:

  • the identification of arrangements and other ‘modernising’ adaptations of chorales in terms of text, melody and setting;

  • the search for quotation forms of smaller and larger specific song sections or entire melodies within non‑liturgical music;

  • an analysis of allusions and chorale‑like topoi in relation to harmonic progressions and individual cadential turns, for instance;

  • an examination of the forms of representation and notation as well as the paratexts, with a view to aesthetic evaluation in new aesthetic and genre‑related contexts.

This project has begun to bring together what might become an extensive digital corpus for hymn settings in a standardised form. Such a resource would confer great opportunities for the systematic investigation of the aesthetic implications of this cultural heritage and its later usage. Beyond the immediate gain in knowledge, this would extend to far‑reaching corners of the art music tradition. Collaboration among researchers with different methodological perspectives working on this repertoire on a global scale – from data sciences and digital musicology to music theory and music history – not only allows us to expect the creation of large datasets but also to organise the multiple perspective and multimodal questions with regard to this corpus of seemingly ‘simple’ hymn settings.

Acknowledgements

Even with a multi‑author paper, there are almost always (many) others involved. We thank all such contributors and list some individual in CONTRIBUTORS.txt files on the relevant repositories.

Funding information

For this article, we acknowledge financial support by DFG within the funding program Open Access‑Publikationskosten. This research was also supported by the Swiss National Science Foundation within the project ‘Distant Listening – The Development of Harmony over Three Centuries (1700–2000)’ (grant no. 182811).

Competing Interests

The authors have no competing interests to declare.

Authors’ Contributions

Corpus preparation was distributed by teams as discussed in the main text. The CAU team prepared the Rein, Zinck, Kittel (four‑voice chorales) and Apel data sets; Gotham and Phan made the multiple‑bass corpora (as reported in Remeš et al. (2025)); and Gotham created the Goudimel set. Hentschel devised and curated the data infrastructure and compiled the Bach data.

This paper was written collaboratively with those teams and individuals listed above leading on their respective topics. Beyond this, Matthias and Kathrin Kirsch lead on the historical context in Section 6, and the status of Gerhardt as first author and Gotham as last intentionally reflects their roles.

Notes

[1] In this paper ‘multi‑voice’ refers to music for more than one voice (soprano, ...) and/or instrument. We prefer this to ‘multi‑part’, which may refer to sections.

[2] For a broader discussion of the situation regarding the singing of hymns in German‑speaking regions of the Lutheran confessions c. 1550–1750, see Garbe et al. (2016).

[3] See comments in the previous endnote.

[4] Future plans include the addition of Speer (1692) and Haßler (1608).

[5] See Ju et al. (2017, 2020b), though note that recent advances have required engagement with a wider range of textures (Nápoles López et al., 2021).

[6] On this and the separation of phrases, see the Hauptstimme dataset, (https://doi.org/10.5281/zenodo.15425748), for which a scholarly report is forthcoming.

[7] Composer‑first, repertoire‑second collections (e.g., ‘Mozart piano sonatas’) are more common, presumably partly because that is a project which can be finished. The Open Score collections provide examples of the repertoire‑first mentality.

[10] MusicBrainz (https://musicbrainz.org/) is the largest available collaborative authority file for music recordings whose description level goes down to the individual movement (of works) and track (of recordings).

[11] It is notable here that some ensembles have included movement timestamps (‘chapters’) in their YouTube videos. This is the case, for instance, with the Netherlands Bach Society’s recording of the St. John Passion, BWV 245. Here is a direct link to the start of the first chorale: https://www.youtube.com/watch?v=zMf9XDQBAaI&t=783s (last accessed on 1 July 2026).

[12] This task is a work in progress and may never be fully resolved in a complete and unambiguous way.

[14] For a verbatim transcription of that explanatory text, see the repo’s README https://github.com/Chorale-Corpus/Goudimel_C.

[16] These were evident and discussed already in the 18th century. For example, see the correspondence between G. Ph. Telemann and C. H. Graun, as discussed in Synofzik (2007, p. 337–353).

[18] For instance, Daniel Gottob Türk made a suggestion concerning full stops in combinations with figures in his treatise on Figured Bass. See Türk (1800, p. 49).

[19] The basso continuo gives rise to a broad scope of scholarly questions, that sometimes are based on the so‑called ‘perceived truth’. For example, the figuring‑density in choral‑settings seems to be higher towards the end of a verse line (approaching the fermata). On this and related matters, see Gerhardt and Kirsch (2024, p. 310–312).

[20] We use music21; other options include humdrum and LaTeX/lilypond.

[21] We speak here of general advantages to this method of encoding early music. For other projects adopting similar approaches, see Dumitrescu (2001).

[22] Possible solutions include user‑friendly interfaces for working directly with strings and seeing the results. Existing options like annotation and viewing text‑based applications for text‑based score files include the ‘verovio.humdrum viewer’. These are promising and fall somewhat ‘in between’.

[23] Information on individual dataset creation processes can be found in the corresponding README for each sub‑corpus.

[24] Since some figures currently disappear when exporting to .mei format using MuseScore (see note 15), the .mei files are incomplete in this respect. They can currently only be used for working with notes and metadata. There is a note to this effect in the README file.

[28] The format is extensible, hence not ‘readily’.

[29] The repository (https://github.com/Chorale-Corpus/Goudimel_C) provides simple code for handling this search and the README shows the structure.

[30] Gotham et al. (2023) is the paper on this topic, and here is a direct link to the relevant part of the repo: https://github.com/MarkGotham/When-in-Rome/tree/master/Corpus/Early_Choral/Goudimel%2C_Claude/Psalmes.

[31] Already in 1919, Arnold Schering discussed this ‘reception‑problem’ of early Bach‑choral‑editions, which seemed to be dedicated mainly to purposes of study (Schering, 1919). A telling contemporaneous statement on the ‘meaning’ of Bach‑Chorales stems from a report, written by Peter Grönland (1761–1825). On this topic, see Kirsch (2023).

[32] For more information on this, please refer to the corpus README at https://github.com/Chorale-Corpus/Rein_JB/tree/main/1755_VierstimmigChoralbuch.

[33] For an overview of music‑centric multimodal data, see Gotham et al. (2025).

[34] This appears in many chorale books, including G. Bronner, Musicalisch‑Choral‑Buch, (Hamburg 1721), with the slightly different spelling Es ist das Heyl uns kommen her.

DOI: https://doi.org/10.5334/tismir.226 | Journal eISSN: 2514-3298
Language: English
Page range: 440 - 455
Submitted on: Sep 9, 2024
Accepted on: Jan 31, 2026
Published on: Aug 4, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Kirsten Gerhardt, Johannes Hentschel, Victor Duy Phan, Matthias Kirsch, Kathrin Kirsch, Mark R. H. Gotham, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.