1. Introduction
Musical compositions from large parts of music history have come down to us in written form only—the ability to record audio reliably dates back to the late 19th century. While there are verbal (and thus ‘indirect’) witnesses, such as concert reports about music predating this technology, the best sources to learn about this music are written scores. Accordingly, musicology has adapted methods from other philologies to investigate and publish these documents in ways that support not only musical performance but also scholarly investigation.
Music notation can be considered a visual encoding of music. As a written language, it is typically used to provide performance instructions facilitating the (re‑)production of sound. As a means of communication between composer or scribe and performer, it is important to consider the conditions of that interaction: Who is writing for whom, what is the purpose of the written document, and to what degree may the scribe rely on a shared understanding of the notational conventions of their time on the audience’s side? A document that is intended as an engraver’s copy for a first edition must fulfill different requirements in terms of legibility and clarity than a sketchbook that serves a composer as an aide memoire while trying out differing ideas. While the first should be accessible to other readers (including modern editors), the latter may refuse compliance with common rules of music notation—the notation is not intended to be read or understood by anyone other than the composer. Obviously, this poses special challenges to a critical edition of such a document.
These challenges often relate to incompleteness: A given sketch may not necessarily provide a clef, time signature, and/or meter signature. It might be appropriate to supply them as an editorial intervention, or they may have been omitted intentionally, as the composer had no statement about one or more of these categories in mind. For instance, a sketch such as the one shown in Figure 1 may only seek to explore a sequence of pitches, without making any statement about the duration of these pitches. In this case, the absence of a meter signature is not to be considered as a deficit—the sketch perfectly serves its purpose, and the different expectations of an audience not addressed by the scribe are not inherent to the document itself. Similarly, a note with a solid notehead and a stem might be used to denote a quarter note—or might make no statement about the note’s duration at all.

Figure 1
A sequence of pitches with no durations implied. Ludwig van Beethoven, D‑BNba, HCB BSk 21/69, 1r.
A scholarly edition of music sketches poses considerably different questions than editions of written music material targeting external audiences. When it comes to sketches, the terminology of music—or, more specifically, the concepts implied by simple labels such as ‘quarter notes,’ is often considered inappropriate. A more robust basis for editorial work can be achieved by retreating to the description of what can be seen—a filled notehead with a stem instead of a quarter. This is a paradigm shift in music philology, made necessary by a different type of music document.
Requiring a different language to adequately describe the visual appearance of a given sketch immediately affects the encoding of such documents using markup languages such as the Music Encoding Initiative’s (MEI) framework (see Hankinson, 2011). When music notation itself is a visual code to represent music, music encoding is a verbal or text‑based code. It relies on a defined terminology to describe individual phenomena. If that terminology itself is inappropriate for describing a given feature, using it by means of a markup language results in tag abuse, as described by the Text Encoding Initiative:
The correspondence between a tag X and the semantic function assigned to it by these Guidelines may not be changed; such changes are known as tag abuse and strongly discouraged. (TEI Consortium, 2025, Chapter iv.1.)
For encoding music sketches, this means that common encoding patterns for scholarly editions based on MEI are not applicable.
Beethovens Werkstatt is a 16‑year research project funded by the German Academy of Sciences and Literature Mainz, which seeks to investigate the requirements for and the potentials of genetic editions of music, using methods of digital music philology. This is mostly conceptual work, resulting in prototypes and proofs of concept. For this type of work, Beethoven is an ideal field of exploration, as plenty of scholarly editions and other musicological research on his music is readily available, as are high‑resolution scans of relevant documents. This facilitates focusing on the methodological research that is at the core of Beethovens Werkstatt (even though it would be highly desirable to apply these new methods to a more diverse set of composers). Currently, the project is in the final stages of a work package exploring methods for how to properly edit sketchbooks in digital form, taking into consideration all the conceptual difficulties described above. While this is ongoing work, and other (composers’) sketchbooks are likely to require modifications to these concepts, this article will introduce the major data modeling challenges faced during this research, including tools and workflow considerations. Where other projects like Mazzagufo et al. (2024) concentrate on the combination of verbal and musical texts to trace the compositional process across different source types, our focus is on improving the handling of visual information in MEI and increasing the transparency of editorial interpretation and intervention that inform a scholarly edition. As these considerations are deeply integrated into a complex data model covering different perspectives on sketchbooks, it seems necessary to discuss this data model in its entirety.
1.1 On diplomatic transcriptions
Sketchbooks by Beethoven and other composers have been subject to research and scholarly editions for decades. Important publications include Dagmar Weise’s edition of the Grasnick 3 sketchbook (Weise, 1957) or Sieghard Brandenburg’s edition of the Keßler sketchbook (Brandenburg, 1978). The approaches taken by these two editions differ significantly. Weise tried to preserve the general layout of the Grasnick 3 sketchbook, as shown in Figure 2.

Figure 2
D‑B Mus. ms. autogr. Beethoven Grasnick 3, as edited by Weise (1957).
This approach makes it very easy to align the edition and the original document. It improves the legibility of the notation by using a standardized music font. However, it does not provide additional space: areas with dense notation are equally dense in the edition. Perhaps even more relevant is that Weise’s edition does not provide sufficient guidance to separate different sketches on a page and understand the sometimes confusing writing order(s) to enable a reader to differentiate and analyze individual sketches. In essence, her approach is a purely visual reproduction of the page, with very little assistance provided in accessing the musical text found on that page.
Brandenburg, in his edition of the Keßler sketchbook, took a significantly different approach, as can be seen in Figure 3.

Figure 3
A‑Wgm: A 34 (Keßler sketchbook), as edited by Brandenburg (1978).
While Weise prioritized the document, Brandenburg focused on the music text of individual sketches. Any revisions are provided in the ossia systems above the text, or, in other words, they are aligned according to their semantic position in the text—not to their physical location in the document. Not being bound by the original layout of the document, Brandenburg has room to provide additional context, including occasional hints on the physical location. However, despite these hints, his edition is not easy to align with a facsimile of the document.
In Feder (1987, p. 138), two different types of diplomatic editions are introduced. The first is labeled as ‘faksimileartig’ (‘facsimile‑like’), while the second is called ‘in normaler Typographie’ (‘using standardized typography’). Feder provides very little explanation for the facsimile‑like diplomatic editions. According to him, the aim of such an edition is ‘a facsimile‑like reproduction of the correctly read text of the source (including its author’s corrections).’ This means that the main contribution of this edition is that the text it provides represents the editor’s reading, based on their experience with the author’s habits, while still mimicking the overall layout of the source. For the second type, Feder states that the intention is a ‘faithful, but typographically and, regarding its layout, normalized reproduction of the correctly read text of the source, including composer revisions, provided in any format, with custom line and page breaks.’ Here, the focus is on the text, and the original layout isn’t as relevant. It is quite interesting that both types, almost ideally represented by the respective editions of Weise and Brandenburg, are subsumed under the same term diplomatic edition. Obviously, this term is broad enough to allow for quite different approaches. Both take different perspectives and have mutually exclusive benefits— they cannot substitute for each other. Beethovens Werkstatt seeks to integrate both approaches, aiming to better bridge the gap between document and text.
1.2 On Notirungsbuch K
The sketchbook used to explore these concepts is called ‘Notirungsbuch K’—this name was assigned to it shortly after Beethoven’s death. Used in early 1823, it mostly contains sketches for the first movement of the Ninth Symphony op.125 and late sketches and revisions for the Diabelli Variations op.120. A small number of sketches can be assigned to other works, and, for about one fifth of the ~480 sketches contained, no connection to a finished work by Beethoven could be identified so far.
Today, Notirungsbuch K is not accessible as such; it has been split into several parts. There are seven known fragments, some of which are now bound together with other materials. These fragments are located in three libraries: Beethoven‑Haus Bonn (D‑BNba), the Berlin State Library (D‑B), and the Bibliothèque nationale in Paris (F‑Pn). Together, all seven make up 84 of the presumed 88 original pages. More details about Notirungsbuch K, its original binding order, the remaining fragments, and the music contained can be found in Brandenburg (1984), Cox (2021), Johnson et al. (1985), and at https://beethovens-werkstatt.de/zum-notirungsbuch-k/.
2. Reconstructing Manuscripts
2.1 IIIF considerations
Before even considering the musical challenges of sketchbooks, the first necessary step in making Notirungsbuch K accessible to modern users is to provide a reconstruction of the original manuscript. Conceptually, the International Image Interoperability Framework (IIIF) has solved this issue. A IIIF ‘manifest usually describes how to present a single compound object such as a book, a statue or a music album.’ Meanwhile, a IIIF canvas is ‘a virtual container that represents a particular view of the object’ (Appleby et al., 2020b, Chapter 2.1). Manifests contain sequences of canvases. To break this technical terminology down, manifests represent documents, and canvases represent individual pages in these documents. However, this relationship is captured through explicit linking—a manifest points to the canvases it holds, in the order in which they are contained. The URLs used for pointing make no assumption about a base URL or shared folder holding all the page images. Instead, each page may be addressed independently of others by an absolute URL. Grossly simplified, IIIF allows for a list of individual links to the scans of the pages of a document to be collated within a single resource.
This concept can be easily adopted in an MEI file. Here, a <facsimile> element represents a given document, and <surface> elements contained therein represent the pages of said document (The Music Encoding Initiative, 2023). Each <surface> may then contain a <graphic> pointing to digital scans. As with IIIF, these pointers may reference scans hosted by different libraries at their respective web servers, so that it is possible to use an MEI file to provide the reconstructed original page order of Notirungsbuch K by putting the links to all seven fragments into the correct order.
2.2 Reconstructions versus current documents
Following this approach raises the important question of which document or sequence of pages should be encoded: the reconstructed Notirungsbuch K, or the current documents in which it has been transmitted. Today, at least for the fragment included in the so‑called Landsberg 8/1 sketchbook (one of about eight sketchbooks named after 19th‑century collector Ludwig Landsberg; see Johnson et al., 1985), the order of pages is different from when Beethoven used those pages. Traditionally, an editor would need to take a decision here and present their reconstructed document order, almost certainly leaning toward the old order of Notirungsbuch K. However, there might be different proposals for reconstructing the order of pages. In fact, Brandenburg (1984) proposed a slightly different document configuration. While being out of scope for Beethovens Werkstatt, it should be possible to trace the full history of Notirungsbuch K from a single manuscript to the seven fragments known today, with all known states in between. Therefore, our model needs to address both the current order of pages in modern documents and one or more historical reconstructions. However, we only have access to scans of modern documents, not the former Notirungsbuch K. Accordingly, the model should reflect that the latter is compiled from modern scans, which can be put in different contexts.
Based on these arguments, solely using MEI’s <facsimile> and its child <surface> elements to reconstruct the sequence of pages does not seem appropriate, and Beethovens Werkstatt decided to support this obvious approach with additional markup. For this, the project relies on the use of <foliaDesc>, with its child <folium> and <bifolium> elements. These elements capture the collation of the manuscript, as documented in The Music Encoding Initiative (2023, Chapter 3.7.1.5). In essence, <folium> and <bifolium> use attributes that point to the (inner and outer) recto and verso faces of the sheet they represent. The targets of these pointers are <surface> elements—so the solution is a combination of <facsimile> and <foliaDesc>. Every modern document is captured in a separate MEI file using both <facsimile>, which describes the sequence of available scans, and <foliaDesc>, which represents the binding of the individual leaves. For Notirungsbuch K, the project uses another MEI file, which contains no <facsimile> but does make use of <foliaDesc> to reconstruct the original page order. In this case, the pointers from <folium> and <bifolium> do not target <surface>s in the Notirungsbuch K reconstruction file but instead the ones in modern documents. This separation of concerns supports any number of competing reconstructions, while also permitting access to capture the fragments in their current context.
2.3 Coordinating coordinates
As mentioned earlier, our reconstruction of Notirungsbuch K is based on fragments transmitted in multiple modern documents. Luckily, scans of all these documents are available. However, these scans differ, and, in this context, the differences in resolution are inevitably relevant. To be able to display (historically) facing pages from different modern documents side by side with correct dimensions, it is necessary to use measurements other than the pixel image dimensions. Obviously, the document’s real‑world dimensions are ideally suited for this.
Another benefit of using both <surface>/<graphic> and <folium>/<bifolium> is therefore the ability to provide two different measurements for each page. While the @width and @height attributes on the first are used to provide pixel dimensions of the available images, the dimensions given in <foliaDesc> are given in millimeters—they describe the document instead. Both dimensions are necessary and cannot be calculated from one another. However, even with both image and page dimensions provided, it is still not possible to calculate a pixel‑to‑millimeter conversion factor, as basically every scan includes a margin of unknown dimensions.
In order to address this issue, Beethovens Werkstatt uses Media Fragments, which are a W3C Recommendation that support specifying a region of interest as an additional URL parameter (see https://www.w3.org/TR/media-frags/#naming-space). Such a media fragment, describing the position of the page within the image, is appended to the link to the IIIF Image API info.json (image metadata) file of the available scan. As seen in Listing 1, this link is not required to end on ‘…/info.json’— the content of this file is served from the ‘…/page5.jpg’ address instead. A second URL parameter is used to indicate the rotation of the page within that image, which ranges from about −1° to about +1° for the available scans of Notirungsbuch K. Not considering rotation probably would have only a minor effect on most pages, but since all later measurements in our data rely on the accuracy of the conversion between pixels and millimeters, it seems highly desirable to get this as correct as possible.

Listing 1
<surface> element with two <graphic> children. The first points to an IIIF image and uses a Media Fragment plus a custom ‘rotate’ parameter to specify the location of the document page within the scanned image; the second refers to an SVG file containing all shapes on that page.
An alternative and equally valid approach for this that has been discussed extensively would be the use of the corresponding features of the IIIF Image API (Appleby et al., 2020a). With this, it is equally possible to select a (rectangular) region of interest within an image, and it also supports rotation. So, with two technical options available, what are the arguments to not rely completely on IIIF and instead combine it with both media fragments and an additional custom URL parameter?
One argument is that, in IIIF, rotation is applied only after selecting a rectangle. While, in our custom solution, we use the same method, for an image requested using the IIIF Image API, this has consequences we were not willing to accept. An IIIF‑only link will return an image file that contains just the requested rectangle, to which a rotation will be applied. This rotation results in a larger bounding box rectangle with ‘empty’ corners—if the image server supports rotation in finer resolution than 90° steps, which some libraries relevant for us do not. This means that rotation isn’t safe to use and needs to be dealt with client‑side anyway. In that case, using the IIIF Image API alone seems to be misleading, and so we decided to use a software design based on additional media fragments instead.
With rotation addressed, it is now possible to determine a conversion factor between real‑world units (measured in millimeters) and image units (measured in pixels). In our data model for diplomatic transcriptions, we mostly rely on using real‑world units—most everything is stored in millimeters, as this is independent of the scan resolution of different pages and leads to more consistent values across the different fragments of Notirungsbuch K. In addition, it makes the data model more compatible with values often provided in printed editions’ source descriptions, such as page dimensions and system heights. Knowing the conversion factor, it becomes possible to automatically extract these and other dimensions from pixel data, as mentioned below. With coordinates resolved, it becomes possible to approach the contents of the sketchbook.
3. Writing Zones
Since its start in 2014, Beethovens Werkstatt has been tracing Beethoven’s pen strokes on all manuscript pages that are subject to our case studies (Cox et al., 2016, p. 18). This means that every character and every stroke drawn on the page is copied using a tablet and saved as an SVG <path> element. These <path> elements make it possible to interact with the facsimile: rendered as an invisible layer on top of the page, it becomes possible to click on and highlight features such as writing zones (WZs) or individual notes. This allows one to unambiguously address every written symbol without any verbal description and base all following musicological arguments on Beethoven’s handwriting.
While it is certainly tedious work to generate these SVG shapes, requiring about 8–10 hours of work for average manuscript pages, we expect that AI may significantly speed up the generation of these SVG shapes in the future. While optical music recognition may produce somewhat similar data in an intermediate step, the focus here is really on tracing pencil strokes, not on giving any interpretation of what the resulting signs might represent. Or, to speak with de Saussure (1916), this step is about signifiers but not (yet) the signified. The concept of agnostic transcriptions, as coined by Iñesta (2019), Ríos‑Vila (2021), and others takes a similar direction, but with an important distinction: While Iñesta (2019, p.14) ‘approach[es] it as a machine translation problem’ to go from these agnostic transcriptions to what they refer to as semantic transcriptions, we consider this a significant step in the editorial process which, at the time of writing, was without sufficient experience (or, in other words, ground truth) to be able to automate this step, at least for Beethoven’s sketches. However, once sufficient data have been produced using FX, it may certainly serve for training purposes. Until then, the current manual approach itself is valuable in the hermeneutic editorial process for how it formalizes an approximation to the manuscript content and helps to grow an interpretation.
In our SVG files, the shapes are grouped by WZs. WZs are what we consider areas of notation that are coherent and establish a context. While a thorough definition for the project’s glossary is still work in progress, for the purposes of this paper, it is fine to consider them as a sketch. There is one important conceptual addition to that, in that WZs are always bound to a page and may not span across multiple pages in our data model. The reason for that is the missing pages in Notirungsbuch K—we simply don’t know if any of the sketches on the available pages continue on one of these. Obviously, the data model needs to be able to represent sketches spanning across multiple pages, but this will be addressed at a later stage.
4. Diplomatic and Annotated Transcriptions
According to our model, each WZ is then diplomatically transcribed. The SVG shapes are purely graphical data and only provide groupings of WZs. The diplomatic transcripts (DTs) are still operating in what is referred to as the visual domain (The Music Encoding Initiative, 2023, Chapter 1.3.1). This means they solely describe the appearance of the manuscript, trying to avoid musical interpretation as much as possible, as this interpretation is the role of our annotated transcriptions (ATs), which operate much more in what MEI considers the logical domain. While this description seems abstract and incomprehensible at first, the distinction becomes very clear when comparing the encoding of a note in both types of transcriptions.
The note encoded in Listing 3 makes use of the most common and basic encoding patterns in MEI. It has pitch name and octave attributes to specify pitch, along with a duration attribute. Together, they give the semantic information that this is a quarter note G4. The note in Listing 2 uses less‑common attributes to achieve something very similar. Here, we have a note that is located on the second staff line (from the bottom) and uses the filled notehead typically associated with quarter notes (and other notes of shorter duration). It also has an upward stem. While such an encoding of notes may look unfamiliar to most users of MEI, these attributes are actually part of the standard, along with many others left out here for clarity. The only exception is an attribute that indicates the (visible) number of flags on a note—for this, a custom @bw:stem.flags attribute has been added and will be proposed for inclusion in the standard at a later point.

Listing 2
A note as encoded in a diplomatic transcription.

Listing 3
A note as encoded in an annotated transcription, which uses regular MEI.
With the information available in both encodings, one would commonly expect both notes to look the same. However, the note in Listing 3 requires additional context for display: Only with information about a clef is it possible to render this note correctly. The note in Listing 2 can safely be rendered visually, even without knowing about the correct clef. Such a clef is necessary, though, to decide about the pitch of a note following this model, and to render it to audio accordingly.
Listings 2 and 3 illustrate the distinction between DTs and ATs. DTs describe what is seen on the page, using musical terminology, but try to avoid providing context as much as possible. With an additional @facs attribute, the notes and other music events here refer to the SVG shapes that lay the foundation for this interpretation—yes, a DT adds a layer of interpretation. It decides about the category of a symbol—for instance, whether it’s a stem or a barline. It also aligns these symbols with the basic ‘layout grid’ of the page it lives on—which staff it belongs to, and the vertical position on that staff.
An AT adds yet another layer of interpretation. Here, musical context is supplied by the editors, such as which clef applies to each staff. Also, notes with unclear duration, as the ones shown in Figure 1, are interpreted and assigned a duration each, and pitches that are not placed ‘correctly’ from a musical perspective are corrected—Beethoven’s sketch notes sometimes seem to be misplaced by a second. Verbal text is normalized, so that an ‘allo’ written by Beethoven resolves into ‘allegro,’ and so on.
Another important aspect of ATs is that they are explicitly not bound to single pages—an AT may cover a single WZ, and thus a single DT, or it may put multiple ones into context and read them as a sequence of WZs that are conceptually one sketch. Where the focus of our DTs is the writing space, ATs concentrate on context and ‘the music’ that can be found in and performed from these sketches. This means that ATs need to follow references like ‘Vi = de’ signs that indicate the continuation of the writing in a different place, much like a dal segno in regular notation. These ‘navigational signs,’ which are quite common for Beethoven and his contemporaries, make use of the Latin imperative vide, meaning look. They are used mostly in the case of corrections or other interruptions of a musical text, where the logical next measure (or any other unit) is not immediately following in regular reading order but is instead written elsewhere for lack of space. The start of this pointer is typically indicated by the syllable ‘Vi = ,’ and the target is identified by the syllable ‘= de.’ However, as these connections are not always unambiguous, this introduces another level of interpretation in ATs.
In summary, an AT involves the editorial interpretation and intervention necessary to read the sketch as actually performable music. While it is not possible to render a DT into sound, an AT can be sonified or performed and thus listened to. However, ATs are still sketches. While they may occasionally include information about instrumentation, they are still not meant for direct performance. Accordingly, Beethovens Werkstatt will support a rather technical sonification using MIDI but not expressive performances in order to avoid giving wrong impressions—these sketches are an aide‑mémoire for the composer, and turning them into actual music is part of later stages in the compositional process.
Coming back to the traditional concepts of diplomatic editions by Weise and Brandenburg introduced earlier, our concept of DTs is closer to Weise, while our ATs are closer to Brandenburg. However, as every step contains a certain level of interpretation, it becomes clear that neither can claim to encode the document or the text exclusively—those categories don’t seem to be binary opposites but instead describe the abstract ends of a continuum connecting the visual and logical domains in music (The Music Encoding Initiative, 2023, Chapter 1.3.1). Our DTs and ATs just take two positions in this spectrum, with the intention to cover as much of it as possible, either directly or through implicit information, which can be derived from combinations of DTs and ATs.
To make these combinations possible, it is crucial to keep connections between the different files. Where a DT uses @facs to refer to the underlying SVG shapes already, an AT uses @corresp attributes to point to the corresponding DT elements. That way, even an AT is indirectly founded on the facsimile—for every note, it is possible to see how it was transcribed diplomatically and how it was written down by Beethoven. Obviously, it is equally possible to operationalize these links in reverse order—for every shape, it is possible to retrieve both DTs and ATs. This means that users can freely navigate through an edition based on this model and can always access adjacent and complementary perspectives on the dataset. Where the agnostic transcriptions by Iñesta (2019) are just an intermediate step, stored in a pragmatic internal format before moving to the ultimately desired semantic transcription, DTs and ATs are considered equally important aspects of one shared data model, focusing on and documenting different steps in the editorial process and both using MEI as a framework to implement this model.
5. Managing Workflows—The Facsimile Explorer
The stack of encodings between facsimile and musical text—from SVG shapes over DTs to ATs—gradually moves from the perspective of document to a perspective of text. A significant challenge is how to generate data conforming to this model efficiently. To assist data preparation and management as much as possible, Beethovens Werkstatt is developing a tool called Facsimile Explorer (FX), which already supports large parts of the project’s data model introduced here. Without FX and the automation provided by other related workflows, it would not be possible to maintain and operate a data model of this complexity. Unlike Plaksin (2023), FX is not aiming to explore novel user interfaces for digital scholarly editions but addresses the editor’s perspective only. At the same time, the visualizations developed for it have sparked the project’s imagination and will certainly influence future user perspectives on the data.
5.1 Software architecture
FX is a web application that runs in the editor’s browser. It is available through a Docker image (see https://ghcr.io/beethovenswerkstatt/facsimile-explorer). Written in JavaScript, it makes heavy use of the GitHub REST API (see https://docs.github.com/en/rest) and loads XML data from the project’s data repository at https://github.com/beethovenswerkstatt/data. These data are loaded into the browser as document object model (DOM), which is quite efficient considering the file sizes loaded and offers sufficient performance for the Notirungsbuch K data. Changes are then made in‑browser to the DOM representations of these files, and editors may commit these changes back to GitHub from within FX. Accordingly, the app requires user authentication using GitHub accounts. While it would be highly desirable to make this configurable in the future, currently, only users with write permissions to the Beethovens Werkstatt GitHub account will get access to FX. The README.md file at https://github.com/BeethovensWerkstatt/facsimile-explorer/ documents how to deploy a custom version of FX using other datasets, though. However, as all interaction with the Git repository makes use of the GitHub API, it is currently not possible to use other Git services. A common API for services such as GitHub and GitLab would certainly help overcome this limitation—for now, the project needs to prioritize the use of its limited resources, even though the current solution is certainly not ideal.
When changes are made through its user interface, FX will allow the editor to commit these changes through a dedicated dialogue (see Figure 4), which lists all files changes, and pre‑generate a commit message that can be adjusted as necessary.

Figure 4
Commit dialog in Facsimile Explorer.
As multiple editors from the project team may work simultaneously, it is not trivial to prepare the commits for these changes. If no other commit has been made since the last refresh of the editor’s client, a direct commit is possible. If this is not possible because of other interfering commits, the changes made locally will be pushed to an auto‑generated, new remote branch. In most cases, FX will then be able to merge this new branch automatically to the base branch. In this case, the new branch will be deleted by FX to keep the repository clean. The editor will get a response that informs them that the situation could be automatically resolved. If, however, there are real conflicts, and an automatic merging is not possible, the remote branch will be kept, and the editor will receive a different message that will point to the conflicting branch on GitHub, from where it needs to be resolved manually. This way, no work will ever be lost. However, since changes are made on the level of individual elements, such conflicts are very rare.
The development of FX is following a participatory approach, where developers and editors closely collaborate and jointly design, implement, review, and improve the tool. This has led to pragmatic and effective decisions in the software architecture, which is shared in large parts with the Cartographer App (see https://github.com/Edirom/cartographer-app, https://cartographer-app.zenmem.de/). While some of these architectural decisions (such as depending on the GitHub API) seem less than ideal in hindsight, they are the outcome of an iterative process that includes the exploration and evaluation of necessary or wanted features along the way, and the restricted development resources of Beethovens Werkstatt allow only limited optimizations of the tools. However, FX and other tools developed by the project are not considered production‑ready software but rather as proofs of concept, illustrating the potential of the project’s data and concept models.
5.2 Setting up documents
The first step in data preparation for our data model is to set up the document files—both for the modern documents holding the fragments of Notirungsbuch K and for the file holding the reconstruction itself. As this is a one‑time operation for each document, no special functionality was built into FX. Instead, Cartographer App was used, as it supports opening an existing IIIF manifest file and generating an MEI file with a corresponding sequence of <surface> elements pointing to the same scans from it. This process worked for the modern documents, as all seven fragments are held by libraries that provide IIIF manifests. Generating the MEI file for the reconstructed Notirungsbuch K required manual compilation from the corresponding segments of these seven fragments, taking into account the original page order in the context of Notirungsbuch K.
5.3 Handling pages
On the first tab of FX (see Figure 5), the editor may adjust pixel and millimeter dimensions of the currently selected page, helping to clarify the relation between the scanned image and the document page, as described earlier. With these data available, it is possible to add a millimeter grid as overlay to the page, allowing for measurement of distances between features on the page. For the future, it is planned to add a ruler tool that allows for measurement of distances in arbitrary angles.

Figure 5
Page tab within Facsimile Explorer.
In addition, the editor is supposed to enter the position and dimensions of all staves on that page. Based on some default assumptions about margins, FX proposes dimensions for the first staff entered, which can then be adjusted as necessary. The second staff will then copy the dimensions of the first and will be placed at a reasonable distance below. From the third staff onward, dimensions and distance will be based on the values entered for previous staves. Each staff can be adjusted as necessary, as the second‑but‑last staff on Figure 5 shows, where Beethoven extended the staff lines manually.
These data are stored in a custom new element called <rastrum>, which is added to the existing <layout> element.1 A rastrum is both the tool used to write (one or more sets of) five staff lines onto paper (in German: Rastral) and the resulting pattern of these staff lines (in German: Rastrierung). Each page that is included in the reconstructed Notirungsbuch K will use the existing @decls attribute on <surface> to point to a unique <layout> element. In the case of machined paper, it would be possible to have the same layout of rastrums on multiple pages, but this is not the case for most manuscripts.
Coming back to FX’s user interface, the left side of the Page tab shows a list of all pages of the currently selected document (besides Notirungsbuch K, all modern documents can be chosen as the current context for display). While similar lists are available on all tabs of FX, the additional details provided for each page vary. In this view, FX indicates whether an SVG file with all shapes on the current page is properly linked in the data or not, whether the page dimensions within the image have been set, and how many staves have been entered for the current page.
Another important aspect of the page list on the left side is the double numbering seen there. The first numbers on the far left are the page numbers in the reconstructed Notirungsbuch K. Following that, the siglum (identifier) of the document that holds this page is given, together with the page number within this document. This facilitates orientation in the data model, as all DTs and ATs are organized by the modern documents to permit creating multiple reconstructions, either competing alternative interpretations or ones covering different times.
5.4 Dealing with Writing Zones
The next tab of FX is dedicated to organizing WZs. It depends on an SVG file with all shapes written to the current page, as mentioned earlier.
When these shapes are traced using commercial tools, the editors try to reproduce the original pen strokes—everything that was written with one stroke of the pen will be represented by one SVG <path> element, even when the ink may have faded out in the middle of the stroke, as sometimes occurs on larger symbols such as slurs. If, in contrast, a music symbol is written with multiple pen strokes, such as a note with a separate notehead and stem, these strokes are supposed to be captured independently. This tracing strategy results in the greatest flexibility and helps minimize later revisions to the shapes, which are possible but bring a certain risk of invalidating references because of changing element IDs.
The main purpose of this tab of FX is to organize the shapes on the page using the different available WZs. In the underlying data, this means that the corresponding shapes are moved into an SVG <g> (group) element that represents this WZ. To do this, the editor just needs to select the intended WZ from the list on the right and then click on all shapes that are considered part of this WZ. The same applies to writing layers, which are used to capture different genetic stages (mostly revisions) within a single WZ, whenever it is possible to clearly identify them.
The list on the right side of Figure 6 not only shows all available WZs (and contained layers) on the current page but also gives a small visual preview that indicates where on the page this WZ is to be found. These data are generated from the bounding box of all contained shapes and are stored in an MEI <zone> element in the corresponding source document. With every change to the shapes contained, this element is automatically updated. The main benefit of this <zone> is improved performance, as such bounding boxes are also relevant for other tools using different software stacks (mostly XSLT/XQuery), which are significantly less efficient to calculate such positional data on the fly.

Figure 6
Writing zones tab within Facsimile Explorer.
The source documents contain not only this <zone> element but also the necessary structures for the WZs themselves, since they are more than just the SVG <g> elements. These structures depend on MEI’s <genDesc> and <genState> elements. These are nested into four levels, as shown in Listing 5.

Listing 4
The use of the custom <rastrum> element inside the <layout> element already provided by MEI. All dimensions are given in millimeters, and relative to the page, with an origin on the top left corner.

Listing 5
Nesting of <genDesc> and <genState> elements inside a document file, representing writing layer 1 of WZ 1 on page 1 of the current document. @xml:id attributes have been omitted, and references in @corresp attributes are adjusted for better legibility.
There are three levels of <genDesc>, each having an @ordered="false" attribute, which means that the encoding order of child elements does not imply a genetic order during Beethoven’s writing process. The outermost <genDesc> is used to capture genetic states at the document level. It contains one <genDesc> for each page of the document. This connection is established by pointing to the <surface> element’s @xml:id. Every WZ is primarily represented by a third‑level <genDesc>, which has a @label displayed in the FX user interface, and points to the corresponding SVG <g> element. This same structure is repeated for each writing layer in that WZ by using a <genState> element. Having this structure for all ~480 WZs on all 84 pages of Notirungsbuch K allows to draw connections between them as they are found, using attributes from MEI’s att.linking class (see https://music-encoding.org/guidelines/v5/attribute-classes/att.linking.html). This would allow a complete genetic edition of the sketchbook—which is clearly beyond the scope of Beethovens Werkstatt and almost certainly impossible to reconstruct within a reasonable timeframe anyway. Instead, the project seeks to explore each WZ independent from each other and focuses on their final state, without attempting to reconstruct potential earlier versions except for some sample cases that will help to illustrate the potentials of this model.
5.5 Annotated transcripts
The next tab of the FX user interface deals with ATs. Conceptually, the next step of interpretation following SVG shapes and WZs would be DTs instead, with ATs following only after that. While this is true for the concept model, a much simpler and faster workflow is possible by addressing ATs first. If DTs were generated based on SVG shapes only, significantly more information would have to be provided and entered by the editor, and two steps of linking would be necessary. The workflow described here and in the following section speeds up the generation of DTs considerably.
In our data model, ATs are quite regular MEI files that are mostly compatible with the MEI Basic profile, which matches the typical needs of score‑writing applications. MEI Basic is intended for interchange with other formats and was introduced by Hankinson (2024). Our AT files are generated using common score‑writing applications and receive only modest adjustments afterward, mostly concerning metadata and a stricter use of elements within <staffDef>. While ATs are not intended to faithfully reproduce the layout of a WZ, they do contain <pb> (page beginning) and <sb> (system beginning) elements and also preserve visual details like stem directions. At the same time, they normalize spelling of words (‘Allegro’ instead of ‘Allo:’). Tools like MEI‑Friend (see Goebl and Weigl, 2024); https://mei-friend.mdw.ac.at/) are used to enter and proofread these data.
When no AT is available for a given WZ yet (see WZ 3 on page 2 in Figure 7), FX allows importing of a new AT file, which will then be integrated into the data model. The imported file will be stored at a defined file path and displayed in FX. However, as an AT may span multiple WZs, FX also allows one to add a WZ to an existing AT that started at a different WZ. The AT shown in Figure 7 covers a total of five WZs, as seen in the side panel on the right. These assignments are captured in the @target attribute of <source> within the AT file, which points to one or more <genDesc> at the WZ level in the document file (see line 3 in Listing 5). In addition, the <pb> mentioned earlier points to the corresponding <surface> in the same file using @corresp. To not conflate different types of information, an additional <annot> is placed next to <pb>, which points to the same <genDesc> as above. This is necessary for having multiple WZs on the same page, being part of a single AT.

Figure 7
Annotated transcripts tab within Facsimile Explorer.
The AT itself is rendered by Verovio (see https://verovio.org and Pugin 2014) and placed in the center of the screen. After rendering, the SVG generated by Verovio is slightly modified, adding boxes on top of the continuous system that indicate the current page or WZ and staves where the transcribed music is found. The information about the occupied staves is entered when importing a new AT file for a WZ. At this point, a suggestion is made based on the overlap of SVG shapes belonging to the WZ and the bounding boxes of the rastrums found on the page, which the editor may adjust as necessary. No other work is necessary on ATs within the FX.
5.6 Diplomatic transcripts
The main purpose of FX is the last tab, dealing with DTs. These may be initialized for any WZs that already have an AT. When initialized, a new MEI file with a predetermined file path is generated. This file uses a highly customized version of MEI, as introduced above. Like ATs, DTs use @target on <source> to reference the <genDesc> of the WZ they transcribe, but, following the model, they will always reference only one.

Listing 6
Basic structure of a diplomatic transcript.
As the content model is significantly different than MEI’s regular <score>, a custom element <draft> is introduced. In Beethovens Werkstatt, it is supposed to encode a single WZ, but other projects might use it for page‑spanning sketches, etc. A <drafts> container allows for multiple <draft> elements in one MEI file. Again, once these new elements have proven their value, they will be proposed for inclusion in the official MEI schema.
The content model requires a <pb>, which indicates on which page the transcribed WZ starts. A custom <system> element captures one system or accolade of music—the set of staves that are supposed to be read together. Since it is quite common for Beethoven’s sketches to change the number of staves from accolade to accolade, each <system> has a <scoreDef> to specify the staves used. This is done by referencing the <rastrum> in the document file. Subsequent <system>s may occupy any combination of staves on the page, as Beethoven often uses any free space on the page to continue writing—be it the immediately following staves or somewhere further away, either above or below.
Each <system> also contains a <section>, which is used as container for <staff> elements. As our DTs are purely visual, a <measure> element would be inappropriate. If a barline is identified as such by the editor, it may be transcribed using the <barLine> element, which has no structural implications in MEI. While it is possible to have multiple <layer> elements, we always use only one and avoid this type of interpretation in the context of DTs.
Figure 8 shows the user interface of FX for the Diplomatic Transcripts tab. On the top, there is a facsimile pane that highlights all shapes belonging to the currently selected WZ either in green (already transcribed as a DT element) or blue (yet to be transcribed).

Figure 8
Diplomatic Transcripts tab within Facsimile Explorer.
The second pane from the top shows the AT of the currently selected WZ. Here, all components that are transcribed in the current DT are also highlighted in green. After having transcribed all shapes, available in the document, this means that all components left uncolored in the AT must be considered as editorial supplements. In a future revision, the Liquifier Tool (see below) will allow to export an adjusted version of the AT encoding that explicitly wraps all these components in <supplied> elements, even though the actual data model holds this information only implicitly.
The third pane is again a facsimile‑based view, but this time with an overlaid DT rendered from the data transcribed. Here, it is possible to click on components to open them in the fourth pane, an XML editor. Currently selected in Figure 8 is a tie in the second staff. However, since pitches are not determined in a DT, it is not possible to distinguish between slurs and ties. Accordingly, we encode both as <curve> elements, which share the same model as both these more specific elements. For selected (control) events, it is possible to directly change values per drag and drop in the third pane, like the @bezier attribute on the selected <curve>.
New elements are transcribed by clicking on an SVG shape in the top pane and then on the element from the AT that is based on this shape. Doing so, FX knows what type of element needs to be transcribed and where in the DT file this needs to be inserted. Then, all semantic information is stripped and converted to purely visual data—a G4 in treble clef will translate into a notehead on the second line from the bottom, and so on. The horizontal position is calculated from the SVG shape, relative to the top left corner of the <rastrum> the element is placed on, properly considering rotation. While the new DT element will receive a link to the SVG shape, the AT element will receive a link to the DT element. The DT element may then be adjusted as necessary using either the visual controls provided on some of these elements or the XML editor. This workflow ensures reasonable default values for DT elements and thus significantly speeds up the preparation of DTs, which are the most complex part of the data model.
5.7 The liquifier tool
The Liquifier Tool is a small NodeJS‑based Docker app that generates certain assets automatically on commit to our data repository (see https://github.com/BeethovensWerkstatt/liquifier) and adds them into a cache folder in the data repository. These assets include enriched versions of our MEI files, like ATs with added <supplied> elements (see above) or ATs encoded compatible with the MEI Basic profile, as documented by Hankinson (2024). Such files are easier to reuse for other use cases, like analytical or MIR tasks trying to compare the sketches to the final works. Accordingly, the Liquifier will export additional MIDI versions of the AT, facilitating reuse of the project data beyond the provision of ground truth for specialized Sheet Music Information Retrieval (see Ríos‑Vila 2024), as mentioned earlier. However, for direct project use, it will also render both DTs and ATs for any changed WZ into SVG files. Since these files will be rendered the same way for every user in our final edition, this offers a significant boost in performance, as it is not even necessary to load the Verovio JavaScript library into the client’s memory.
While this improved speed is certainly nice to have, it might not be necessary if Beethovens Werkstatt would stop here. The main purpose of Liquifier is something more ambitious, though. Calculating fluid transcriptions (FTs) (see below) is considerably more expensive on processing than rendering regular DTs or ATs, so the main motivation for this little tool is facilitating the generation of FTs. It is currently work in progress.
6. On Fluid Transcriptions
DTs are rendered based on the document dimensions. When rendering ATs of the same WZ, we bring them to the same scale, so that DTs and ATs make use of the same coordinate space. This lays the foundation for a highly innovative concept of FTs, as first mentioned in Kepper et al. (2026). FTs seek to morph between DTs and ATs, which significantly improves legibility of very densely notated sketches by disentangling them slowly. Figure 9 shows the endpoints of FTs. The exact sequence of intermediate steps between these endpoints is still under consideration; they include the normalization of spacing and shapes and the addition of supplied elements such as clefs and time signatures. Technically, this happens by comparing the relative positions of corresponding elements in DTs and ATs and taking the differences between these values as input to an SVG transformation. As long as both transcriptions use the same SVG structures and dimensions, applying these transformations is almost trivial but requires the rendered versions of both DTs and ATs as input.

Figure 9
Transformation endpoints of a fluid transcription.
One conceptual challenge of FTs is that they are not able to resolve multiple competing genetic states—if there is variation, only one version may be transformed into a normalized AT at a time. This can be addressed by requiring the user to select a genetic variant to be rendered. This will adjust both the AT and DT endpoints of the FT, for which the transition can then be controlled using a simple slider control. While proofs of concept for FTs exist already, a proper implementation and integration into the digital editions of Beethovens Werkstatt is still work in progress.
7. Conclusion
Capturing minute visual details of music documents such as sketchbooks requires a data model of significant complexity. However, with the ability to render these details into diplomatic transcriptions and careful connections to both the individual graphemes captured as SVG and the normalized appearance of more traditional transcriptions, it becomes possible to ‘liquify’ the notation and increase the accessibility of densely notated sketches considerably. While traditional approaches of diplomatic transcriptions like the ones by Weise and Brandenburg take static positions within the continuum that connects the philological concepts of document and text, the model we propose explores the possibilities of combining multiple such positions into an integrated perspective that focusses on the connections between these categories, not their differences. This is a major paradigm shift in (music) philology and will have significant impact in the field.
While we can already foresee the potentials of this new concept, further research is required. Consistent modeling can only be achieved by investigating the best way to diplomatically encode different music symbols. Similarly, it requires more experiments to explore the limitations of what can and should be rendered by Verovio as the de‑facto default rendering engine for MEI and what needs pre‑ and postprocessing and custom rendering.
Finding answers to these questions will take time. What is clear already, though, are the costs for preparing data of this complexity. While the amount of work needed is still considerable, the use of FX makes it possible to transcribe larger manuscripts like Notirungsbuch K according to the model proposed herein, which enables interactive transcriptions that help in bridging the gap between document and text.
Competing Interests
The author has no competing interest to declare.
Note
[1] As with other custom elements and attributes, <rastrum> will be proposed for addition to the official MEI schema once enough experience has been gathered. For now, the development of these proposals can be traced online at https://github.com/BeethovensWerkstatt/data/tree/dev/odd.
