Skip to main content
Have a personal or library account? Click to login
Publishing and Performing Digital Musicology on the Web: Signature Sound Vienna Cover

Publishing and Performing Digital Musicology on the Web: Signature Sound Vienna

Open Access
|Jul 2026

Full Article

1 Introduction

Musicologists have been understood as a core audience for music information retrieval (MIR) since the field’s earliest pre‑history (Kassler, 1966), and musicologists’ needs have motivated MIR development since the first International Society for Music Information Retrieval (ISMIR) conference (Bonardi, 2000). Yet, the impact of MIR technologies on musicology has remained limited (Borsan et al., 2023).

We report on ‘Signature Sound Vienna’ (SSV), an interdisciplinary digital musicology research project investigating the Vienna Philharmonic Orchestra’s (VPO) New Year’s Concert (NYC) series by digitising, studying and disseminating findings pertaining to performance recordings, scores and other evidence objects associated with this famous concert series.

Beyond our immediate musicological interests, our project aims to contribute toward making digital approaches accessible to wider scholarly audiences. Integrating domain expertise in music informatics, performance science, historical musicology and semantic Web technologies, we have compiled, developed and interconnected digital tooling into workflows, facilitating the creation of FAIR‑data musicological corpora. We have also produced interoperable Web applications for conducting and disseminating digital musicology through an iterative process of development and immediate feedback undertaken in close collaboration between technologists and music scholars. These applications include the mei‑friend editor and mark‑up tool for music encodings, the Listen Here! environment for tool‑assisted close listening to score‑aligned audio collections and the Primal Platform for review and interaction with music annotation Linked Data. These tools are loosely interconnected through shared Linked Data structures employing W3C standards and established Semantic Web ontologies. The applications are browser‑based, with no further installation requirements. Their Web‑based architecture provides scope for future development, re‑using and further enriching the underlying data. The co‑creative development process between informatics and musicology specialists reduces friction in integrating the software within our own musicological investigations and aims to provide useful tooling for broader application in music scholarship.

In this paper, we provide an overview of our digital musicology tools and workflows. We contextualise our work in its paradigmatic and technological background (Section 2), describe semi‑automated workflows to digitise and publish Linked Data characterising our collected performance recordings as a FAIR data corpus (Section 3), detail our tools supporting musicological inquiry and dissemination of findings in connection with supporting evidence objects (Section 4), offer a case study of their application in musicological scholarship (Section 5) and discuss broader implications (Section 6), before concluding and outlining plans for future work in Section 7.

2 Background

2.1 Musicologists’ attitudes to digital methods

ISMIR papers have repeatedly identified an understanding of user contexts, information needs and behaviours as a prerequisite to developing useful MIR systems (Lee and Cunningham, 2013; Schedl et al., 2013). Reflecting on the first 10 years of ISMIR (2000–2009), Downie et al. (2009) called for actively encouraging user participation, including by musicologists, to help create ‘truly useful music‑IR systems’ (p. 17). Though digital resources are important to contemporary musicologists (Inskip and Wiering, 2015; Laplante and Sauvé, 2023), scholars are heterogeneous in self‑assessed technological expertise, report struggles in integrating software within their research practices and may even perceive digital approaches as a threat to traditional musicology. Tools supporting traditional methods are more readily accepted than those conducting distant forms of analysis on the scholar’s behalf – though digital capacities for grounding assertions in empirical data are understood to strengthen the validity of findings (Laplante and Sauvé, 2023, p. 37).

Consequently, MIR technologies are rarely used in musicology: Analysing the 1,055 ISMIR papers published during 2012–2021, Borsan et al. (2023) found only 28 reporting on datasets, methodologies, code and/or tools whose use is subsequently reported in the musicology literature, even when surveying interdisciplinary venues such as the International Conference on Digital Libraries for Musicology and the Journal of New Music Research.

2.2 FAIR approaches to music research

Provenance is an important notion in both traditional and digital scholarship, (i) placing evidence objects (e.g. performance recordings; scores) within historical, cultural, social and pragmatic contexts, particularly in terms of ownership and possession, and (ii) modelling scholarly assertions and analytical processes and their outcomes in terms of agency, parameterisation and evidence basis. Correspondingly, contextual information on agents, materials and contexts has been incorporated into musical and musicological data models using the Linked Data paradigm, informed by FAIR (‘Findable, Accessible, Interoperable and Reusable’) principles of research‑data management (Wilkinson et al., 2016). Originating in the natural sciences, these principles have received growing interest in music – see, e.g., Dix et al. (2022), Hofmann et al. (2021) and Weigl et al. (2021) – and more broadly in the digital humanities (e.g., Hawkins, 2022).

The OMRAS2 project (Cannam et al., 2010) pioneered the use of Linked Data approaches to managing music information. This was achieved in part through the creation of a family of Semantic Web vocabularies and ontologies centred on the Music Ontology (Raimond and Sandler, 2012), which has been applied in research but also in broadcast media institutions. The Transforming Musicology (Crawford et al., 2017) and FAST (Sandler et al., 2019) projects respectively investigated the application of such approaches in music scholarship and in music industry contexts. AcousticBrainz (Porter et al., 2015) and CALMA (Bechhofer et al., 2017; Page et al., 2017) processed decentralised, user‑contributed information, publishing large collections of user‑provided recordings and music‑analytical metadata as semantically‑enriched Web corpora. Dunya (Porter et al., 2013; Serra, 2014) established a system for exploring multicultural music research collections, assembled as part of the CompMusic project (Serra, 2014), by means of culturally‑relevant contextual relationships. The TROMPA project investigated the connection of such repositories over the Web, allowing their data to drive task‑specific interfaces catering to diverse user communities, including music performers and enthusiasts alongside scholars (Weigl et al., 2019, 2021). Polifonia took further steps in this direction, establishing new shared ontologies and tools for data integration, transformation and unified access to music heritage collections (de Berardinis et al., 2023). The ongoing LinkedMusic project aims to lower barriers of access to music search, harnessing large language models to provide internationalised natural language querying across a large, heterogeneous collection of music datasets (Pond et al., 2025).

The Music Encoding Initiative’s MEI format (Hankinson et al., 2011) has seen widespread application as a machine‑readable, music‑semantic frame of reference in many of these projects. MEI provides a means of encoding music notation documents, typically rendered as visual scores using the Verovio toolkit (Pugin et al., 2014). MEI documents are readily parsable by machine agents. By producing visualisations in the scalable vector graphics (SVG) format retaining the MEI encoding’s (XML) element hierarchy and identifiers, Verovio facilitates implementing browser‑based applications centred upon interactive digital scores (Pugin, 2018).

MEI’s musically‑meaningful XML model affords granular integration with Linked Data (Weigl and Page, 2017), which may flexibly target music elements using their XML identifiers in media‑fragment uniform resource identifiers (URIs) (Troncy et al., 2012) or via music‑specific selection schemes (Viglianti, 2016). MEI and Linked Data technologies have been combined in a number of research projects connecting scores and music recordings with extra‑musical context, including in modelling annotations and scholarly claims (Lewis et al., 2019, 2022; Sanfilippo et al., 2025), in associating musical meaning with timed media (Lewis et al., 2019) and in dissemination (Dreyfus et al., 2025; Lewis et al., 2018). A complementary approach is taken by Dezrann (Ballester et al., 2025), a platform for the annotation and synchronization of music data, which integrates (rather than interconnects) disparate multimodal datasets by converting them to custom unified data representations.

3 Establishing a FAIR Research Data Corpus

SSV is digitally capturing, analysing and interpreting a wealth of music information relating to the NYC series, while facilitating reuse and reinterpretation of these data by the wider digital music research community. Our workflows and tools support investigations spanning the history of the NYC series, applying varied but complementary technological approaches in music scholarship. We have assembled a data corpus pertaining to these concerts, including scores, recordings and feature data characterising the recorded signal, performance and catalogue metadata and information drawn from the broader historical context. These are interconnected with relevant authority records and published using interoperable vocabularies, adhering to FAIR principles. Projects overly focusing on data management may delve deeply into data and library science but risk losing their music research focus (Jensenius, 2021). Accordingly, our team comprises researchers in historical musicology and performance studies, alongside information science, and our workflows incorporate parallel processes of music informatics and music scholarship.

3.1 Workflow for corpus creation

We have acquired and digitised source materials on physical media (CDs, DVDs and LPs) relevant to the NYC series (Weigl et al., 2022). Their acquisition is non‑trivial: only NYC recordings from 1987 onward remain available for purchase. We thus searched second‑hand shops and flea markets to expand our corpus, incorporating other orchestras’ recordings of frequently‑performed NYC compositions by consulting Bielefelder Katalog Klassik and other catalogues. Our collection includes more than 80 albums comprising audio from 61 New Year’s Concerts (43 ‘complete’ concert recordings, alongside compilations) and 30 albums by other orchestras performing relevant repertoire.

In digitising our physical media collection of audio recordings, we are informed by the approach taken by Dunya (Porter et al., 2013) to manage large collections of recorded audio and associated metadata. We captured our tracks as uncompressed WAV files. As in Dunya, MusicBrainz is used to store and organise the editorial metadata obtained from covers and booklets of our physical media. Our tracks were identified using the open‑source AcoustID audio‑fingerprinting service1 and labelled with discographic metadata obtained from MusicBrainz using the Picard Tagger (Stutzbach, 2011). Labels were validated against (physical) liner notes. Corrections or previously missing records were contributed back to AcoustID and MusicBrainz where required. Validated metadata were converted to readily‑processable format using Picard’s ‘Generate Cuesheet’ plugin (one cuesheet per recording), including structured information on both album and track levels. A custom cueToRdf Python script processed these sheets, extracted audio features and generated Linked Data descriptors to characterise our collection. The script generates RDF graphs adhering to the Music Ontology, with distinct sub‑graphs listing predicates and objects for all subjects (entities) of the following types:

  • mo:Record, a music record (CD, LP, DVD, etc);

  • mo:Track, a track on a record;

  • mo:Signal, the audio signal captured on a track;

  • mo:Release, a record as a commercial product (e.g. with liner notes and a catalogue number); and

  • mo:Performance, a (recorded) musical performance by one or more agents or groups of agents (e.g. a conductor, an orchestra, a soloist).

An instance‑level graph is illustrated in Figure 1.

Figure 1

Simplified entity‑level graph describing the 2017 NYC rendition of An der schönen blauen Donau. Rectangles: entity labels. Rounded rectangles: Web resources.

Entity URIs may be dereferenced to access their corresponding data in our corpus. MusicBrainz URIs are provided where relevant.

In Dunya, digitised and semantically enriched music data are stored within a set of relational and graph databases. To minimise technical requirements and promote sustainability and workflow reusability beyond our project scope, we instead store our Linked Data graphs as static files (Section 3.2). An additional union graph was produced to facilitate cross‑dataset operations (e.g. SPARQL queries) and to provide a convenient representation for archival in versioned research data repositories such as Zenodo.

Serialisations are provided in different RDF formats, allowing machine‑agents to specify their preferred content type for processing (Section 3.2.1). Audio‑derived features are published following a non‑consumptive approach to reproducible research on copyright‑restricted materials (Organisciak and Downie, 2021). Waveform envelopes are computed to provide accessible visualisations of recordings for public dissemination (Section 4.4). Industry authority metadata (e.g. International Standard Recording Code) are provided to facilitate copyright‑conformant access by project‑external researchers.

3.2 Publishing our Linked Data corpus

w3id.org is a URL redirection service operated by the W3C Permanent Identifier Community Group providing identifier namespaces – in our case, https://w3id.org/ssv/ – for the formulation of proxy URIs pointing to a resource’s hosted location. Content‑negotiation allows machine agents to specify their preferred RDF serialisation format through standards‑compliant specification of acceptable content type(s). This service facilitates interoperability and supports long‑term sustainability by allowing the dataset to be moved at a future date without impacting (potentially project‑external) Linked Data referencing its resources, following established best practices for implementing FAIR Linked Data on the Web (Garijo and Poveda‑Villalón, 2020).

Our data is stored on GitHub. Hosting is provided free of charge, with expectations for long‑term storage. Cross‑origin resource sharing compliance permits Web applications to directly load and process our resources. All files are versioned, with access flexibly provided through predictable ‘raw’ GitHub user‑content URIs. The data can thus be exposed in a specified state according to publication context by incorporating a label corresponding to a particular branch, tag, release or commit number. Versioned data can be explicitly exposed using w3id.org proxy URIs.

3.2.1 Example: organising track‑level linked data

The following example is provided to clarify our approach to organising Linked Data representations as flat file collections for flexible, sustainable, versioned data management. The RDF representation for track 20 on the VPO’s 2017 NYC performance recording – An der schönen blauen Donau – in its latest version (on the default ‘main’ branch), in RDF‑XML serialisation, is stored at: https://raw.githubusercontent.com/Signature-Sound-Vienna/data/main/data/track/2017.rdf#202

An equivalent JSON‑LD serialisation is stored at: https://raw.githubusercontent.com/Signature-Sound-Vienna/data/main/data/track/2017.jsonld#20

While the default branch houses the most up‑to‑date version of our data, it is desirable to be able to maintain particular versioned states for the sake of reproducibility. A TISMIR branch, maintained to support the arguments in this paper, is stored at: https://raw.githubusercontent.com/Signature-Sound-Vienna/data/TISMIR/data/track/2017.ttl#20

The equivalent proxy URI for this data in its latest‑updated state, regardless of serialisation, is: https://w3id.org/ssv/data/track/2017#20

And, in the version maintained for this paper: https://w3id.org/ssv/TISMIR/data/track/2017#20

These versioned proxy URIs are used to reference entities within the dataset’s graphs, ensuring internal consistency. Notably the small set of project‑specific RDF predicates maintained in the https://w3id.org/ssv/vocab# namespace are not versioned within the graphs. This is done in order to retain cross‑version compatibility (Garijo and Poveda‑Villalón, 2020).

Content negotiation allows data consumers to request the RDF serialisation most convenient for their purposes. To demonstrate, the following curl command retrieves the data in the Turtle (.ttl) format: $ curl ‑H “Accept: text/turtle” ‑L https://w3id.org/ssv/TISMIR/data/track/2017#20

4 Tools for Digital Music Research and Dissemination

We have produced interoperable, Web‑native tools to conduct and disseminate musicological research on our FAIR Linked Data corpus, developed collaboratively between technologists and music scholars and refined in response to stakeholder feedback in musicology, music encoding, music informatics and digital libraries.

Our central musicological questions address the extent to which a ‘signature sound’ can be discerned across decades of VPO NYC performances, conductors and interpretations, along with how such patterns may be situated within broader historical and institutional contexts. While our tools are developed to help address these specific research questions, they aim to be reusable by the wider digital musicology community.

Traditional close‑reading and close‑listening methods yield rich musicological insight but become less feasible as the scale of a collection increases. The typical Digital Humanities response is to pivot to notions of ‘distant reading’ (Moretti, 2013) – and, equivalently, ‘distant listening’ (Have and Enevoldsen, 2021) – relegating the immediate processing of (digitised) material sources to algorithms operating over feature data. Indeed, we extracted such content‑derived features during construction of our dataset and are preparing publications on distant‑listening analyses.

Here, we focus on a complementary approach: informed by musicologists’ typically sceptical attitudes to distant analyses (Section 2.1), we use digital technologies to make close analyses tractable over large collections. We use audio‑to‑audio and audio‑to‑score alignments and assistive user interfaces to support immediate and precise comparison – and cross‑modal annotation – spanning performance recordings and music scores. This reduces the cognitive overhead involved in finding correspondences across recordings.

Alignments are generated using the SyncToolbox Python package3 for memory‑restricted multi‑scale dynamic time‑warping using chroma onset features (Müller et al., 2021). We have developed a script to apply this technique to large collections, aligning many renditions of a piece to a chosen reference recording to project an isometric (20‑ms) reference grid to all audio timelines. In addition, the reference timeline is aligned to a corresponding score encoding. Structural variances among performances, especially in repeat patterns and codas, are catalogued and encoded using MEI’s expansion elements to normalise temporal reference across disparate interpretations. The structurally‑most‑complete performance in each instance serves as the reference rendition.

4.1 Mei‑Friend

Music scores enjoy a privileged position as evidence objects in scholarship and provide a natural reference when comparing collections of recorded renditions. MEI encodings are highly suitable for such purposes (Section 2.2), but manually encoding MEI‑XML is laborious and requires significant technical expertise. We have developed mei‑friend4 (Figure 2) to accommodate: a browser‑based music encoding editor (Goebl and Weigl, 2024) capable of importing and converting music encodings in various formats to MEI using an integrated instance of the Verovio toolkit. Pre‑existing score data obtained through OMR or typeset using MuseScore thus bootstrap the encoding process before conversion to MEI for the ‘last mile’ of correction, validation and enrichment. An adaptable user interface presents a synoptic view of both MEI XML and (Verovio‑generated) graphical scores. Notation elements selected in the visual score are correspondingly selected within the XML editor and vice versa. Graphical menus with keyboard shortcuts support modification or insertion of MEI elements. An integrated GitHub module facilitates collaborative encoding. A facsimile panel permits direct comparison and interlinking with (regions of) scanned source documents. A MIDI player sonifies the score, highlighting notes and following the score during playback. An enrichment panel supports editorial interventions and markup, including creation and display of annotations using the MEI <annot> element. Stand‑off Linked Data structures adhering to the Web Annotation Data Model (Sanderson et al., 2017) and the Music Annotation Ontology (Lewis et al., 2022) may be generated or imported, supporting independent annotation of Web‑hosted MEI files. These annotations are Linked Data structures addressable from external contexts, including other Web applications, for subsequent enrichment and reuse.

Figure 2

A Linked Data annotation (corresponding to the example in Section 5.2.2) is captured on an encoded score using mei‑friend.

4.2 Listen Here!

Listen Here! is a Web application5 for machine‑assisted close‑listening to aligned audio collections (Weigl et al., 2023). Recordings are listed in a side‑panel and are loaded into or removed from a comparison view on user request. Dozens of recordings may be simultaneously visualised as playable waveforms using WaveSurfer.js,6 and load‑on‑demand access is provided for arbitrarily large numbers of recordings. Tempo‑curve visualisations are available as waveform overlays, showing either absolute performance tempo (quarter notes per minute) or relative deviation from the average tempo at a given moment across configurable subsets of the collection.

Audio visualisations are vertically stacked (Figure 3), recalling the interface of the Sonic Lineup desktop application.7 On playback, the sounding (active) waveform is highlighted; a cursor indicates playback position and can be manually reset via mouse‑click. By clicking on another waveform, it is made active and the previous one becomes inactive; playback seamlessly switches to the new recording, using the alignment grid projections to arrive at the corresponding playback position with fine‑grained accuracy. Visual indicators may be set at corresponding time positions across exemplars. Optionally, playback jumps back to the closest marked position when switching between recordings, greatly facilitating close listening to specified moments in the music across renditions.

Figure 3

Listen Here!: Linked Data (Music Annotation Ontology) annotation corresponding to the example in Section 5.2.2 projected onto performance renditions.

Hand in hand with mei‑friend’s support for stand‑off annotations, we have implemented equivalent functionalities for audio recordings within the Listen Here! tool. Linked Data structures targeting MEI elements may be imported and projected onto corresponding (audio‑to‑score‑aligned) timeline positions in each rendition. Annotations are visualised through corresponding coloured regions upon each loaded waveform. Using a Select Recordings button, the Linked Data are enriched with timed media‑fragment URIs referencing timeline intervals on the selected recordings.

4.3 Data‑Level integration

Our tools support decentralised workflows in which Linked Data structures are generated and iteratively enriched using independent Web applications targeting different music modalities. This is achieved by adherence to explicit machine‑interpretable vocabularies and data models – an inherent affordance of the Linked Data paradigm – casting the processing of music information as ‘digital music objects’ experiencing ‘semantic music flows’, accreting data and contextual information along the way (De Roure et al., 2018; Sandler et al., 2019).

Our applications implement this approach using the Solid platform for social Linked Data (Mansour et al., 2016), a decentralised solution supporting contribution and sharing of user‑created content, including for digital music research applications (Weigl et al., 2020). Solid Pods are both identity providers (allowing users to authenticate when logging in to independent applications) and user‑administrated storage spaces supporting selective data sharing across applications and with other users.

The Music Annotation Ontology (MAO; Lewis et al., 2022) provides a framework to capture, extend and refine cross‑modal music annotations. Both mei‑friend and Listen Here! are capable of generating, loading and enriching MAO Linked Data structures (e.g. by adding new MAO Selections to existing Music Objects). Unlike in Sanfilippo et al. (2025), the primary aim is not to model the granular semantic structure of musicological claims but rather to map annotations to their targeted music representations, which may be interrelated across different modalities and varying levels of abstraction.

To facilitate interactions with this model, we apply a storage and retrieval pattern to aid discovery of the relevant data across applications. A user logs in to different Web applications (currently, mei‑friend and Listen Here!, though other applications, e.g. presenting music‑analytical visualisations, would be equally supported) using Solid. On authentication, each application establishes the existence of a music‑annotation data–integration container within the user’s storage, linked from the user’s Solid profile. Similarly, containers for MAO music objects, for Web Annotations pertaining to music data integration and for a resource discovery service are established.

Each application presents Web resources – in our case, MEI encodings and performance recordings – to the user, who is given the option to make a selection upon the presented resources and to identify this selection as the basis of a new music object. This prompts the creation of a chain of Linked Data structures within the MAO music object container, starting with a MAO Selection object that directly targets the selected resource fragments using media‑fragment URIs – here identifying MEI elements or time intervals on audio recordings. The process is repeated with increasing abstraction for MAO Extract and MusicalMaterial – see Lewis et al. (2022). The discovery service container is checked for existing collections referencing a particular Web‑resource (e.g. MEI file, audio recording). These collections are modelled using schema.org, a widely‑used vocabulary for the description of Web resources, as DataCatalogs8; each collection is about9 its respective Web resource. Where such a collection does not yet exist, it is created, and references to each MAO music object or Web Annotation structure are added in connection to the resource being annotated, using the dataset10 property.

This pattern enables independent applications to establish and interact with shared storage locations for mutually‑relevant data resources. Code‑level integration is not required; the interoperability required to enact the ‘music flows’ across applications as described here is achieved through integration on the data. Figure 4 provides a simplified, idealised illustration of such a workflow using our tooling.

Figure 4

Simplified annotation workflow: (i) Evidence objects supporting a finding are formalised as MAO Music Objects – here, using mei‑friend’s tools for score selection and annotation. (ii) Multimodal enrichment of these structures is carried out using compatible tooling – here, relevant aligned timeline intervals are identified using Listen Here! and captured as additional MAO Selections. (iii) A textual annotation (Web Annotation) is associated with the MAO Music Object stack. (iv) The combined annotation structure can be viewed accessibly (to human readers) through a browser, using the Primal viewer, or it may be requested as machine‑readable Linked Data. Note: Discovery service details withheld for simplicity; see illustration in Weigl et al. (2023).

4.4 Disseminating our findings through hypermedia publication

Using our approach, scholarly insight is explicitly connected to constituent evidence objects (score fragments; audio‑recording intervals), allowing findings to be disseminated as hypermedia publications juxtaposing scholarly analysis, annotated digital scores and audiovisual materials. Using Linked Data to model both evidence and scholarly annotations allows each of these entities to be addressed from external contexts, including from more traditional forms of dissemination such as journal articles – see, for example, our long‑form report of an investigation into Johann Strauss II’s Kaiser‑Walzer (VanderHart and Weigl, 2025). In short, this approach supports decentralised scholarly communication, annotation, re‑interpretation and re‑contextualisation in the spirit of Linked Research (Capadisli, 2020).

Music scholars cannot be assumed to possess expertise in Linked Data models, serialisations or other particular technologies, nor can such expertise be required if data adhering to these models are to be made useful to this music‑domain‑expert audience. Appropriately user‑friendly views over the relevant data structures are required.

To realize such human‑readable hypermedia publications, we have developed Primal11 (Figure 5), a browser‑based viewing and interaction tool for music annotation Linked Data. The tool is instantiated with the URI of a Linked Data entity (supplied via the ?obj query parameter); configurably traverses along graph connections originating from this ‘root’ entity, building an internal graph representation as it proceeds; and registers encountered target objects of interest for custom rendering. In its default configuration, the tool assumes the root entity and graph relationships relevant for traversal to adhere to the Web Annotation or MAO data models. Target objects are specified using media fragment URIs of MEI elements and specified intervals of audio recordings, in compliance with the Music Encoding and Linked Data framework (Lewis et al., 2018).

Figure 5

Primal (excerpt): Our Linked Data annotations – https://w3id.org/ssv/TISMIR/annot/oa/ExDonau in the example from Section 5.2.2 – redirect to human‑readable hypermedia publication views when their URIs are retrieved through a Web‑browser. Requests specifying RDF serialisations (e.g. by machine agents) instead retrieve the underlying Linked Data via content negotiation.

Following traversal and subsequent validation, specialised renderers construct human‑readable views of different aspects of the data. These include:

  • descriptive content of scholarly annotations, rendered as text;

  • annotated score fragment(s), rendered to SVG using Verovio, with annotated elements highlighted;

  • annotated audio interval(s), highlighted upon corresponding waveform visualisations12;

  • metadata, including the identity of the annotator and information on the annotated recordings;

  • A navigable graph visualisation showing a configurable subset of the currently‑traversed Linked Data graph, generated using the open‑source Mermaid.js diagramming tool13; and

  • A full data listing of the graph contained at the currently‑navigated RDF resource (highlighted on the visualisation), presented as nicely formatted JSON‑LD.

The different views are presented in sections, navigable using a persistent menu. Sections are intentionally ordered to be welcoming to all music‑interested users, starting with textual, score and audio contents and relegating the more technical metadata, graph visualisation and data listing (JSON‑LD) views. These latter views remain readily discoverable and accessible, allowing users with greater technical affinity, musicologists interested in explicitly engaging with the underlying data and students in pedagogical contexts to gain deeper insight into the underlying corpus and data model.

5 Case Study: Signature Sound Vienna

We now evaluate the utility of the technologies introduced above through our investigation of the VPO’s NYC series within the SSV project.

5.1 The process

SSV operationalises a hybrid methodology combining close and distant listening. Our tools facilitate iterative cycles of exploratory listening, structured annotation and Linked Data–driven analysis across a multimodal corpus of historical and contemporary performances.

Listen Here! enables comparative close listening across aligned audio‑to‑score collections, facilitating the identification of both fine‑grained performance nuances and larger structural differences within and across recordings. We used this tool to closely examine the 10 most frequently performed works in the New Year’s Concert repertoire – all composed by members of the Strauss dynasty (e.g. An der schönen blauen Donau, Radetzky‑Marsch, Kaiserwalzer, Pizzicato‑Polka, and the Fledermaus overture). Multiple renditions of each work were accessed in rapid succession through time‑aligned playback anchored to MEI‑encoded scores. Corresponding historical (public domain) scores were encoded and annotated using mei‑friend, enabling the interactive selection of notated elements, the creation of Linked Data annotation structures with persistent URIs and their reintegration into Listen Here! for targeted analysis and enrichment. This annotation supports a layered analytical process. Listening sessions began with open‑ended, hypothesis‑generating questions – e.g. whether the Vienna Philharmonic exhibits a consistent interpretive profile across decades, or the extent of conductors influence on interpretive contours. Initial subjective observations were subsequently formalised as semantically‑enriched annotations, which served both as analytical anchors and as citable evidence.

Annotations and linked media were contextualized using the Primal platform, which renders the resulting knowledge graph as human‑readable hypermedia. These interconnected evidence objects, accessible via URIs, support transparent scholarly communication, data reuse and further inquiry beyond the scope of the project. This integrated workflow – anchored in traditional close listening but enhanced and rescaled through bespoke computational tools – allowed the project to move beyond anecdotal claims and toward empirically grounded and replicable analysis of interpretive trends, institutional identity, performance style and conductor agency.

5.2 Findings

We conducted close comparative listening across several dozen recordings spanning more than 80 years of New Year’s Concerts and related performances. This analysis yielded insights into evolving performance practices, institutional habits and interpretive norms within the repertoire of the Vienna Philharmonic and beyond. The following, diverse examples illustrate representative analytical outcomes.

5.2.1 Conductor agency/‘same procedure every year?’

In the final bars of the Sphärenklänge waltz (measures 253–258), a striking divergence in dynamic shaping appears across performances by the same orchestra over time: https://w3id.org/ssv/TISMIR/annot/oa/ExSph. In early recordings, a piano marking following an upwards swell into bar 255 is completely ignored by the Philharmoniker; the ensemble maintains dynamic intensity and tempo with minimal attenuation. The reason remains unclear: the marking may have been perceived as musically implausible, absent from the scores used at the time or simply overlooked. It is first with the 1987 performance under Herbert von Karajan that this dynamic shift is fully realized. Karajan introduces a preparatory ritardando in the prior bar; integrates the dynamic shift organically; then reshapes the following passage into a subdued, lyrical reminiscence rather than an emphatic triumph. This seems to have set the tone; all subsequent Vienna Philharmonic recordings use this approach. Notably, Karajan had employed the same interpretation in his 1971 Berlin Philharmonic recording, suggesting that this reading was imported to the VPO rather than organically institutionalized. This case demonstrates how individual, conductor‑specific aesthetics can exert lasting influence, shaping microdynamic architecture within an otherwise stable institutional performance tradition.

5.2.2 Notated versus normative phrasing in An der schönen blauen Donau

A clear divergence between score and performance practice occurs in the opening phrase of An der schönen blauen Donau. In all critical editions, the first flute ascends from E at the end of measure 4 to F‑sharp on the downbeat of measure 5. In performance, however, this ascent is almost universally replaced by a descent to D, aligning the phrase with the melodic contour of the first waltz section to which the introduction alludes: https://w3id.org/ssv/TISMIR/annot/oa/ExDonau. This convention has become so entrenched that performances adhering to the noted ascent in the score now appear anomalous. Only four recordings in our corpus follow the score: three by the K & K Philharmonie under Matthias Georg Kendlinger and one by the New York Philharmonic under Leonard Bernstein (1982). This finding highlights the capacity of tradition to override textual authority, particularly in canonic works. It also illustrates the value of computational listening tools in identifying instances where assumptions about score fidelity warrant reassessment.

5.2.3 Rubato norms in the Fledermaus overture

The Fledermaus overture, especially in the Allegretto passages, contains numerous rubato conventions that are not notated but are deeply ingrained in Viennese performance practice. One representative example occurs in measures 79–83, where a repeated phrase at a fairly steady tempo is followed by a pronounced deceleration that closes the section with a cadential ‘sigh’. While the degree of rubato varies, its expressive shape remains remarkably consistent across Viennese ensembles, including the Vienna Philharmonic, Wiener Volksopern‑Orchester, and Wiener Johann Strauss Orchester under Herbert Siebert: https://w3id.org/ssv/TISMIR/annot/oa/ExFled.

Over time, Vienna Philharmonic performances exhibit a gradual intensification of this deceleration. Outside Vienna, interpretations vary more widely. Some ensembles, such as the London Symphony Orchestra under Charles Mackerras, closely emulate the Viennese model, while others modify it or disregard it entirely. Kurt Schmid, for instance, leading the Philharmonie Lugansk, extends the final pause significantly longer than typical Viennese readings. In contrast, the K&K Philharmoniker under Matthias Georg Kendlinger plays straight through, with minimal temporal flexibility. These contrasts highlight how performance traditions can serve as markers of specific, localised stylistic identity and also how unwritten interpretive habits can, in some cases, supercede notated authority.

5.2.4 Structural variants in Éljen a Magyar!

The quick polka Éljen a Magyar! exemplifies the contingency of structural performance practice. Analysis of the da capo section (bars 111–112), using Listen Here! reveals three competing structural templates across the corpus: a long form with full repeats (https://w3id.org/ssv/TISMIR/annot/oa/ExElj1); a more economical structure popularized in mid‑century concerts (https://w3id.org/ssv/TISMIR/annot/oa/ExElj2); and a more recent, idiosyncratic variation (i.e. Eschwé with the Wiener Johann Strauss Orchester, 2016), which appear informed by a recently published critical edition score by Michael Rot (https://w3id.org/ssv/TISMIR/annot/oa/ExElj3). The decision to play or omit certain sections contributes substantially to the character of the performance and alters audience perception of pacing and symmetry. These structural variants are arbitrary but correlate with available performance editions and conductors’ rehearsal cultures.

5.2.5 Boskovsky‑era performance practices in Vergnügungszug

Measure 94 in another ‘polka schnell’, Vergnügungszug, offers insight into the theatrical performance culture of the Boskovsky era. During Willi Boskovsky’s tenure as concertmaster (1955 and 1979), performances frequently incorporated unscored, irreverent interventions, often associated with percussionist Franz Broschek. These included humorous ‘tuning’ of the Kondukteurhorn (conductor’s horn) (https://w3id.org/ssv/TISMIR/annot/oa/ExVer1); impromptu horn blasts; and the use of props, interruptions and vocalizations.

Such practices, documented in recordings, influenced later international orchestras in performance (https://w3id.org/ssv/TISMIR/annot/oa/ExVer2). This aligns with the performative flair characteristic of Boskovsky’s tenure and reflects a now‑historicized layer of interpretation, one tied not only to a conductor but to a broader performance culture. Such instances exemplify how ephemeral performance traditions may become fixed in recorded history or sometimes transferred or translated between orchestral cultures, serving both as interpretive data and cultural artefacts.

6 Discussion

Our novel digital musicology workflows and tools support the creation of FAIR‑data music corpora and authoring of music‑annotation Linked Data structures using accessible graphical applications, by:

  • digitising collections of recorded media, automatically characterised with Linked Data descriptions in accordance with FAIR research data–management best practices;

  • encoding digital notation in the MEI format;

  • annotating music elements upon such digital score representations, projected onto aligned performance recording timelines;

  • efficiently and tractably conducting tool‑assisted close‑listening analyses through aligned recording collections, even for corpora with many dozens of recorded renditions of a given piece;

  • iteratively constructing citable, machine‑readable musicological Linked Data objects by selectively associating score‑authored annotations and intervals in recordings; and

  • publishing human‑readable hypermedia overviews of such structures in context of their associated evidence objects, for use in pedagogy and scholarly dissemination.

Throughout, we emphasise principles of minimal computing and aim to lower barriers to access of these technologies. The listed capabilities are enacted using open‑source tooling chosen or specifically developed for the described workflows. Decentralised, freely‑provided storage is used both in the working environment and for dissemination: GitHub, the Solid platform for Social Linked Data and the Zenodo repository for research data management. While thousands of open‑source projects depend on GitHub’s free offerings, risks of a change in monetisation strategy or of future withdrawal of access must be acknowledged. Our use of the w3id.org redirection service would then support a transfer to another hosting provider without invalidating our published URIs (assuming the redirection service were to pivot away from its own GitHub use). The Solid ecosystem remains research‑heavy and subject to rapid evolution, though Solid’s open‑source development process emphasises sustainability (Van Herwegen and Verborgh, 2024). While lowering barriers of access, our use of decentralised data storage solutions thus does not obviate requirements for a fundamental understanding of research‑data best practices within the research team.

Our workflows predominantly incorporate user‑friendly tools with graphical UIs, and core musicological activities (tool‑assisted close analyses of scores and recordings across aligned collections; authoring of annotations) are achieved purely from the browser with no further installation requirements. However, several steps of the process still require more technical intervention using command‑line scripts. We have attempted to minimise complexities as far as possible, providing sensible default parameterisations and batching operations such that single invocations operate upon large collections of data. Of course, this in turn carries assumptions: the defaults we provide to the Essentia feature extractor (which we run over the entire corpus), and to the SyncToolbox alignment toolkit (which we run on the digital audio, and on the MEI‑encoded scores, corresponding to our ten focal compositions), will not be appropriate in all music contexts and use‑cases. While mei‑friend facilitates practice and pedagogy around music encoding by exposing key functionalities through graphical menus – providing validation and schema‑informed auto‑completion within the XML editor, and integrating community resources, including the MEI Guidelines, to enable direct documentation lookups for selected elements – its use still requires a foundational understanding of MEI and certainly an openness to learning more. Establishing an appreciation of such issues and their implications on behalf of potential users must be a pedagogical priority in digital musicology, as technological literacy remains a requirement for those conducting research in this space, even where tooling evolves to make music information processes more readily accessible to non‑MIR‑experts.

7 Conclusions, Limitations, and Outlook

We have presented workflows and software tools for conducting and disseminating musicology developed through an iterative interdisciplinary process between music scholars and technologists within SSV, a project investigating the VPO NYC series. The project has aimed to facilitate scholarly use of digital technologies relevant in investigations of large music collections, including scores and performance recordings. It has done so by providing accessible approaches to publish such collections as Linked Data corpora adhering to FAIR principles of research‑data management with minimal computing requirements (foregoing, e.g., availability of institutional repositories); and, by developing a modular collection of Web‑based applications to support scholars in close study and annotation of evidence objects, and in the dissemination of findings.

The cross‑modal ‘integration‑on‑the‑data’ employing independent tools, enabled through the combination of Solid, Linked Data models such as the Music Annotation Ontology, and MEI’s music‑semantic XML structure, feels like a particularly powerful model for future software development in digital music research. It supports complex musicological workflows (complicated from both scholarly and technological perspectives) using comparatively simple applications focusing on specific tasks or modalities. The approach is extensible, allowing new applications – potentially developed by different groups, perhaps with different disciplinary perspectives (e.g. music psychology, music pedagogy, performance science) – to benefit from the capabilities of existing ones, without requiring the installation of additional software by the user. Because the user‑generated, iteratively‑enriched musical and musicological data structures have an application‑independent identity as citable Web resources, they may be directly incorporated into scholarly communication, knowledge dissemination and teaching.

We cannot claim to have rigorously evaluated the usability of the presented tools within the scope of this paper, as the case study presented in Section 5 was conducted in context with the co‑creative, iterative tool‑design process – though mei‑friend has enjoyed widespread uptake by the broader MEI community over the past few years, Listen Here! and Primal are made available as publicly‑hosted Web services for the first time in connection with the publication of this article. Wider user evaluations in pedagogical and research contexts are planned to support future developments. We have recently launched a science communication project dedicated to studying and refining our tools and workflows in the context of disseminating musicology to a wider public audience (Weigl, 2025).

We acknowledge known usability limitations in preprocessing steps – corpus RDF generation (Section 3) and audio‑to‑audio and audio‑to‑score alignment (Section 4) – which currently require command‑line interaction. These steps are modular in our workflow, and we plan to replace them with graphical interfaces (see, e.g. Jiang et al., 2025) and ultimately Web applications in future work.14 However, we feel that our case study already demonstrates the affordances and advantages of the Web‑native, interoperable social Linked Data approach in making digital music research and dissemination more widely accessible, and we hope to have thus contributed in a small way toward bringing MIR and musicology into closer contact in the future.

Acknowledgements

The authors gratefully acknowledge our project colleagues and advisors for their contributions: Markus Grassl, Matthäus Pescoller, Delilah Rammler and Fritz Trümpi. The notion of ‘integrating on the data’ pursued by our tools is informed by longstanding collaborations and discussions with members of the DH group at the Oxford e‑Research Centre, University of Oxford, particularly Kevin R. Page and David Lewis. We thank Anna Plaksin, Laurent Pugin, Thomas Weber, Julia Jaklin, Henning Burghoff and Sophie Stremel for contributing code and documentation to mei‑friend, and Mark Saccomano for advice on our tools’ interactions with the Music Annotation Ontology. We are grateful to the wider Music Encoding Initiative community for valuable user feedback.

Data Accessibility

Signature Sound Vienna’s codebases and datasets are available as follows:

cueToRdf script for Linked Data corpus creation from .cue files (Section 3.1; MIT license): https://github.com/Signature-Sound-Vienna/cueToRdf

Linked Data corpus (Section 3.2; CC BY 4.0): https://github.com/Signature-Sound-Vienna/data

Alignment workflow script built around SyncToolbox (Section 4; MIT): https://github.com/Signature-Sound-Vienna/alignment

mei‑friend codebase (Section 4.1; AGPL 3.0): https://github.com/mei-friend/mei-friend

Listen Here! codebase (Section 4.2; AGPL 3.0): https://github.com/Signature-Sound-Vienna/listenHere

Primal codebase (Section 4.4; AGPL 3.0): https://github.com/iwk-digital/primal

MEI encodings of the 10 most‑performed NYC compositions (Section 5.1; CC BY 4.0): https://github.com/Signature-Sound-Vienna/encodings

Competing Interests

DMW is an elected member of the Board of the Music Encoding Initiative (2023–2028), which is on a voluntary basis. All other authors have no competing interests.

Authors’ Contributions

David M. Weigl: Principal Investigator, Signature Sound Vienna. Co‑developer, mei‑friend. Developer and co‑creator, Listen Here! Developer of Primal, SSV data architecture and project‑specific scripts.

Chanda VanderHart: Project member, Signature Sound Vienna. Co‑creator, Listen Here! Responsible for music scholarship conducted and disseminated within Signature Sound Vienna.

Werner Goebl: Primary project advisor, Signature Sound Vienna. Co‑developer, mei‑friend. Co‑creator, Listen Here! Responsible for MEI encodings created during the project (alongside Matthäus Pescoller and Delilah Rammler).

Funding Information

This research was funded in whole or in part by the Austrian Science Fund (FWF) https://doi.org/10.55776/P34664. For open‑access purposes, the author has applied a CC‑BY public copyright license to any author‑accepted manuscript version arising from this submission.

Notes

[1] Available at https://acoustid.org/; all listed URLs last accessed June 10, 2026.

[2] Note that the #20 fragment completes the URI of that track, but that it is served within a single document corresponding to all tracks of that record.

[4] Available at https://mei-friend.mdw.ac.at. Extensive documentation available at https://mei-friend.github.io

[6] Available at https://wavesurfer.xyz/

[7] A tool in the Sonic Visualiser family of applications; see https://www.sonicvisualiser.org/sonic-lineup/

[11] Available at https://primal.mdw.ac.at

[12] Where copyright restrictions prevent the publication of the corresponding audio, pre‑calculated (non‑playable) waveform envelopes are instead visualised so that an accurate visual stand‑in of the track remains accessible even where sharing audio is prohibited. This is achieved by:

  1. pre‑calculating envelopes (waveform peaks) for every track (mo:Signal) in our dataset;

  2. situating all URIs for copyright‑restricted audio (the vast majority of audio within our corpus) within a https://w3id.org/ssv/audio/ namespace, while hosting accessible audio elsewhere when copyright permits;

  3. instructing http://w3id.org/ssv to return a 403‑Forbidden HTTP response for any /audio/ requests;

  4. instructing the Primal waveform renderer to pivot to the pre‑calculated‑peak visualisation whenever a 403‑response is encountered.

[13] Available at https://mermaid.ai

[14] Since the initial submission of this article, we have implemented a graphical alignment interface within Listen Here! This performs the DTW algorithm for audio‑to‑audio and audio‑to‑score alignment entirely within the user’s browser, with no audio upload required, making it suitable for use with copyrighted materials.

DOI: https://doi.org/10.5334/tismir.315 | Journal eISSN: 2514-3298
Language: English
Page range: 247 - 263
Submitted on: Jun 30, 2025
Accepted on: Mar 30, 2026
Published on: Jul 8, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 David M. Weigl, Chanda VanderHart, Werner Goebl, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.