Skip to main content
Have a personal or library account? Click to login
KRAISLER: A Multi‑Track Dataset of Piano and Violin Duet Recordings for Music Information Retrieval Research Cover

KRAISLER: A Multi‑Track Dataset of Piano and Violin Duet Recordings for Music Information Retrieval Research

Open Access
|Aug 2026

Full Article

1 Introduction

In classical music, the piano plays a central role as a solo instrument, offering a wide pitch range and the capacity for polyphonic expression. Beyond its solo use, it frequently serves as an accompaniment, enriching the performances of melodic instruments such as violin, cello, and flute with harmonic and rhythmic support. In both solo and accompaniment contexts, recordings of expressive piano performances are valuable for a wide range of music information retrieval (MIR) tasks, including automatic music transcription, performance analysis, music source separation, and automatic music accompaniment. This value stems from the ability of such recordings to capture subtle timing variations, dynamic articulations, and the rich acoustic characteristics of real instruments, which synthetic audio often struggles to reproduce with the same depth and nuance.

However, most existing datasets remain limited in scope, particularly when it comes to representing ensemble contexts of Western classical music. Widely used piano datasets such as the MAPS (Emiya et al., 2010), SMD (Müller et al., 2011), and MAESTRO (Hawthorne et al., 2019) datasets consist of only solo performances, often recorded with high‑precision MIDI systems. On the other hand, multi‑instrument datasets such as the URMP (Li et al., 2018) and Bach10 (Duan and Pardo, 2011) datasets provide audio mixtures and isolated tracks for various instruments, but the piano is not included in their instrumentation. Moreover, even when a piano is present, as in datasets like the MusicNet (Thickstun et al., 2017), TRIOS (Fritsch and Plumbley, 2013), or PCD (Özer et al., 2023) datasets, the recordings often rely on score‑to‑audio alignment rather than precise performance capture. This reliance limits their effectiveness for tasks, such as transcription, requiring high temporal accuracy.

To this end, we introduce KRAISLER (‘KAIST Realistic Annotated Instrument‑Specific Live Ensemble Recordings’)1, a new multi‑track dataset of piano–violin duets designed specifically for MIR research tasks in real‑world conditions. It includes synchronized piano MIDI data, separately recorded piano and violin stems processed with varied reverberation, MusicXML scores and their score–MIDI alignment data, note annotations, and beat annotations. The dataset comprises 20 pieces, each excerpted at a duration of approximately 1–2 min. It features high‑quality audio recorded separately by professional musicians in acoustically isolated yet visually connected rooms, enabling clean source separation and controlled performance analysis. The piano performances were played on a Yamaha Disklavier. This MIDI‑controlled piano allows precise real‑time capture of performance data, including note onset and offset times, pitch, velocity, and pedal usage. In addition, we provide three versions of the audio with varying acoustic characteristics: dry, studio, and hall acoustics. The studio and hall versions were generated by adding artificial reverberation to the dry audio. The corresponding musical scores are provided in both MusicXML and PDF formats. These were initially generated using optical music recognition and subsequently corrected for errors by a professional musician. We also provide symbolic alignments between the piano MIDI data and the corresponding score. Violin note annotations were automatically extracted via score‑to‑audio alignment. Using the initial piano beat positions from the piano‑score alignment, researchers with formal musical training manually adjusted the global beat positions to reflect the expressive timing of both instruments to generate precise beat annotations for performance analysis and beat tracking evaluation.

Using these resources, we explore a range of MIR research tasks under realistic acoustic conditions, including piano transcription in both solo and piano–violin duet settings, violin transcription, music source separation, score alignment, and beat tracking. We further demonstrate the dataset’s utility as a benchmark for evaluating model performance.

2 Related Work

In this section, we review publicly available classical music datasets for MIR tasks. Their key attributes are summarized in Table 1.

Table 1

Comparison of publicly available classical music datasets and their attributes for MIR tasks.

InstrumentsDatasetRecorded AudioMulti TrackDirect Piano MIDI CaptureNote AnnotationBeat Annotation
Solo pianoMAPS (Emiya et al., 2010)
SMD (Müller et al., 2011)
MAESTRO (Hawthorne et al., 2019)
ASAP (Foscarin et al., 2020)
Ensemble (no piano)URMP (Li et al., 2018)
Bach10 (Duan and Pardo, 2011)
PHENICX‑Anechoic (Miron et al., 2016)
ChoraleBricks (Balke et al., 2025)
Piano EnsembleRWC‑Classic (Goto et al., 2002)
TRIOS (Fritsch and Plumbley, 2013)
MusicNet (Thickstun et al., 2017)
SWD (Weiß et al., 2021)
PCD (Özer et al., 2023)
KRAISLER (Ours)

2.1 Solo piano datasets

Some datasets have been created for piano transcription research. Among them, the MAPS dataset (Emiya et al., 2010) represents one of the earliest efforts to provide piano recordings with aligned MIDI data for piano music transcription. The recordings have a total duration of about 65 h and were generated either through MIDI‑based software synthesis, simulating diverse acoustic environments such as studio, concert hall, church, and jazz club, or through performances on a Yamaha Disklavier piano. Although it is known to contain annotation errors (Gong et al., 2019), the dataset has been widely adopted in the MIR community as a standard benchmark for evaluating piano transcription algorithms.

The SMD dataset (Müller et al., 2011) is another early contribution in this research area that also employs direct MIDI capture via the Yamaha Disklavier. It includes 50 recordings performed by students from the Hochschule für Musik Saar, each paired with perfectly aligned MIDI files. The dataset features excerpts from Western classical composers, ranging from Bach to Scriabin, totaling approximately 4 h and 43 min of music, with each excerpt lasting around 5–6 min.

The MAESTRO dataset (Hawthorne et al., 2019) marked a transition toward large‑scale data collection and contains approximately 200 h of piano performances across more than 1,000 pieces from nine years of International Piano‑e‑Competition events. It is distinguished by its concert‑quality acoustic grand piano recordings and high‑precision MIDI capture, providing reliable ground‑truth labels for piano transcription.

The ASAP dataset (Foscarin et al., 2020) extends the MAESTRO dataset by pairing each performance with its corresponding musical score with additional annotations; it is further extended in (n)ASAP (Peter et al., 2023) with note‑level alignment.

While the datasets mentioned above provide substantial value for solo piano transcription research, they have inherent limitations in ensemble settings, where the piano serves as an accompanist to other instruments.

2.2 Multi‑instrument datasets

To support music analysis in multi‑instrument settings, various synthetic and real datasets have been developed. While some of these datasets include piano, many datasets focus solely on orchestral instruments, limiting their applicability to piano‑related tasks in ensemble contexts. Moreover, for datasets that rely on score–audio alignment labels, the annotations may lack precise temporal accuracy.

2.2.1 Synthetic datasets

Several synthetic datasets have been developed for multi‑instrument scenarios, offering time‑aligned audio and MIDI rendered with high‑quality virtual instruments.

The EnsembleSet dataset (Sarkar et al., 2022) offers 80 orchestral tracks spanning 6 h synthesized under various simulated recording conditions, with access to both mixtures and isolated sources. The CocoChorales dataset (Wu et al., 2022) consists of 240,000 synthesized chorale‑style tracks using 13 Western classical instruments, amounting to approximately 1,400 h of audio, and is designed for large‑scale harmonic and texture analysis. SynthSOD (Garcia‑Martinez et al., 2025) provides more than 47 h of high‑resolution orchestral mixtures and stems, featuring diverse dynamics, tempo variations, and articulations, making it well‑suited for orchestral modeling.

While synthetic datasets offer the advantage of providing large‑scale multi‑track audio with diverse instrumentation, they remain limited by the absence of real human performances and the fundamental acoustic differences between synthesized audio and real‑world recordings. Moreover, they are confined to orchestral settings, which restricts their suitability for tasks involving piano transcription or ensemble contexts where the piano is present.

2.2.2 Real performance datasets

In contrast to synthetic datasets, which rely on MIDI‑rendered audio, real performance datasets are constructed from recordings of human musical performances captured in natural acoustic conditions.

Some real performance datasets provide multi‑instrument recordings with source‑specific annotations for non‑piano ensemble settings. URMP (Li et al., 2018) provides multitrack classical ensemble recordings spanning a diverse set of instruments, together with score‑to‑audio alignment and pitch annotations. Bach10 (Duan and Pardo, 2011) focuses on four‑part Bach chorales performed by violin, clarinet, saxophone, and bassoon, offering note‑, pitch‑, and score‑alignment annotations. PHENICX‑Anechoic (Miron et al., 2016) and ChoraleBricks (Balke et al., 2025) further provide separated instrumental stems with score‑based annotations, supporting tasks such as score‑informed source separation and transcription.

Other real performance datasets include piano in broader chamber, concerto, and art‑song settings. The RWC Classical Music Database (RWC‑C), a subset of the RWC Music Database (Goto et al., 2002), covers a broad range of classical repertoire, including symphonic, chamber, solo, and vocal works, and provides structural annotations such as themes and repeated sections. For chamber and concerto repertoire, TRIOS (Fritsch and Plumbley, 2013) provides five chamber trio recordings, covering four classical works and one jazz work, recorded in a fully bleed‑free setup with separate instrument tracks and aligned MIDI scores. The MusicNet dataset (Thickstun et al., 2017) provides classical ensemble recordings of 11 instruments including piano, with note‑level annotations for each instrument aligned to MIDI scores. The audio is released only as a single mixture without separate stems, and, due to suboptimal alignment quality, MusicNetEM (Maman and Bermano, 2022) was later introduced with improved annotations. Schubert Winterreise Dataset (SWD) (Weiß et al., 2021) focuses on Schubert’s voice–piano song cycle Winterreise, offering multiple recordings with score‑aligned annotations, including measure, harmonic, structural, and lyrics annotations. The Piano Concerto Dataset (PCD) (Özer et al., 2023) consists of 81 carefully selected 12‑s multitrack excerpts drawn from 15 piano concertos. Each excerpt includes separately recorded piano and orchestral tracks in dry and artificially reverberant versions, accompanied by corresponding musical scores.

Together, these datasets provide valuable resources for transcription, separation, score‑related analysis, and performance studies. However, few datasets combine separate piano audio, high‑resolution piano performance data, synchronized ensemble recordings, and note‑ and beat‑level annotations in live duet settings. A further consideration is temporal coordination between performers. Many existing multi‑track ensemble datasets rely on predetermined tempo guidance, such as conductor videos or metronome clicks, which is well suited to larger ensembles and provides a relatively clear timing reference. By contrast, duet performances typically allow for more flexible coordination, where performers adjust directly to each other’s expressive timing. This makes it challenging to define a shared beat timeline for both performers. KRAISLER complements existing resources by capturing this piano–violin duet interaction together with synchronized audio, MIDI, score‑level annotations, and manually refined global beat annotations.

3 KRAISLER Dataset Construction

3.1 Planning

We recruited one violinist and one pianist from the local area who had professional ensemble experience. The repertoire was selected from pieces the performers frequently performed, so that the pieces could be recorded without requiring extensive additional rehearsal; it comprises violin concertos and sonatas, along with iconic vocal and piano works widely adapted for violin. These pieces cover a broad historical range, from Mozart in 1775 to Kreisler in 1940. From each piece, we selected a segment in which the violin and piano are played simultaneously, focusing on sections where both instruments are active.

Table 2 summarizes the composer and the metadata of the original works, including title, key, and year of composition, along with the duration of each excerpt in our dataset. Note that, since our dataset consists of selected excerpts rather than complete pieces, the key of each excerpt may not necessarily coincide with that of the original work.

Table 2

The list of pieces in the KRAISLER dataset.

No.ComposerTitleKeyYearDuration
01TchaikovskyValse sentimentale, Op. 51 No. 6B minor18821:04
02RachmaninoffPreghiera (arr. by Kreisler from Piano Concerto No. 2, Op. 18)C minor19401:03
03Saint‑SaënsDanse Macabre, Op. 40G minor18741:19
04TchaikovskyViolin Concerto in D major, Op. 35, mov.1D major18781:18
05Saint‑SaënsViolin Concerto No. 3 in B minor, Op. 61, mov.1B minor18801:12
06Schumann3 Romances, Op. 94A minor18491:18
07BrahmsViolin Sonata No. 1 in G major, Op. 78, mov.3G major18791:04
08GriegViolin Sonata No. 3 in C minor, Op. 45, mov.1C minor18871:14
09FranckViolin Sonata in A major, mov.4A major18861:24
10ChopinNocturne No. 20 in C minor, Op. Posth.C minor18301:23
11SchumannDichterliebe, Op. 48, No. 1F minor18401:24
12SchumannDichterliebe, Op. 48, No. 5A major18400:48
13BeethovenViolin Sonata No. 5 in F major, Op. 24, mov.1F major18010:50
14PonceEstrellita (arr. by Heifetz)F major19121:17
15RachmaninoffVocalise, Op. 34, No. 14C minor19121:54
16TchaikovskyMélodie from Souvenir d’un Lieu Cher, Op. 42, No. 3E major18781:28
17MozartViolin Concerto No. 3 in G major, K. 216, mov.1G major17751:19
18KreislerLiebesleid (Love’s sorrow)F minor19051:25
19KreislerLiebesfreud (Love’s joy)D major19051:07
20MendelssohnViolin Concerto in E minor, Op. 64, mov.1E minor18441:09
Total duration25:00

The performance difficulty level was primarily determined using the G. Henle Verlag classification system2, which assigns levels from 1 to 9 for various instrumental works, including piano, violin, cello, flute, and clarinet. Among the works in our dataset, 11 pieces have difficulty levels provided by the Henle website, ranging from 4 to 9, corresponding to medium to difficult levels. It should be noted that these levels reflect the overall difficulty of the original work rather than that of the specific excerpt used in our dataset.

We set an appropriate honorarium and scheduled two recording sessions, each including sound‑checking, technical setup, and a joint rehearsal. All sessions were arranged over two weeks.

3.2 Recording and MIDI capture

The recordings were conducted in two adjacent, acoustically isolated rooms, as shown in Figure 1, separated by a clear, soundproof window that facilitated visual communication between the violinist and the pianist, resulting in a bleed‑free recording setup. Both performers wore wired headphones routed through the control room, enabling real‑time monitoring of each other’s playing to ensure precise synchronization. The violin was miked using a Shure SM57 dynamic microphone, and the piano with a matched stereo pair of large‑diaphragm condenser microphones. Audio was recorded at 44.1 kHz, with 16‑bit resolution using Logic Pro3 with a Focusrite Clarett+ 8Pre audio interface. Each session involved approximately two hours of setup and joint rehearsal, followed by one hour of recording, during which each excerpt was recorded two to three times with brief rest periods inserted as needed.

Figure 1

Recording setup for the dataset: (a) Room 1 contains the Disklavier piano, while (b) Room 2 is used for violin recording, with a soundproof window enabling visual interaction between performers.

Simultaneously, the Yamaha Disklavier’s MIDI output, including note onset and offset events, velocity information, and pedal usage, was recorded into the same Logic Pro project. Although recorded in sync, minor temporal drift was detected after the sessions, and timing was manually corrected during post‑processing.

3.3 Room acoustics and mixing

To simulate realistic acoustic environments beyond the dry recordings, reverb was independently applied to each instrument stem, yielding three versions: dry, studio, and concert hall. We used AudioEase’s Altiverb 84 to apply reverb, using two presets, Paramount Stage for studio acoustics and London Wigmore Hall for concert hall simulation. Considering the wide dynamic range of classical music both within individual pieces and across the repertoire, we avoided a uniform mixing strategy. Instead, equalization, compression, and limiter settings were tailored for each piece. The primary focus was on (1) ensuring balanced volume levels across all pieces, (2) applying noise reduction and compression tailored to each composition’s unique characteristics, and (3) optimizing the reverberated spatial impression of both instruments to maintain a consistent acoustic environment. We performed all mixing procedures under the initial guidance of a professional sound engineer.

Following room acoustics simulation and tailored mixing for each instrument stem, the final mixture was obtained by summing the processed piano and violin tracks. The gain of each track was adjusted so that the piano (signal) and violin (noise) were combined at a predefined SNR of 0 dB. To prevent digital clipping, we applied a limiter to the mixed audio signal and simultaneously adjusted the overall gain.

3.4 Score preparation and note alignment

MusicXML score files offer a significant advantage over sheet images by enabling direct access to symbolic information for score extraction and analysis. To construct the MusicXML dataset, we collected available scores from the MuseScore website and manually created additional MusicXML files from sheet images where necessary. We applied optical music recognition with ScanScore5, a commercial software tool, to produce an initial draft of the MusicXML score, which was later manually corrected by a professional musician. During the conversion of the original sheet music images to MusicXML format, performance recordings were also used as references. Minor discrepancies were left unchanged, whereas substantial deviations or omissions were corrected to better reflect the actual performance. Notes that were clearly not played were removed from the final MusicXML, while small mistakes (e.g., when a pianist played only octave passages without inner voices or when a violin chord was realized as a single melodic line) were retained as in the original score.

Based on the finalized MusicXML files, a note‑wise score‑to‑MIDI alignment was conducted exclusively for the piano parts, using the corresponding piano MusicXML score and the performance MIDI. We used DualDTWNoteMatcher, a DTW‑based score‑to‑performance note matcher and the SOTA model provided by the parangonar library6, to generate .match files. Based on alignment, we automatically annotated each note ID from MusicXML and MIDI, channel, pitch, and onset time.

For violin note annotations, we employed NoteEM (Maman and Bermano, 2022), a framework that simultaneously trains a transcriber and aligns a reference score to its corresponding audio using an Expectation–Maximization scheme. In the E‑step, the unaligned score is warped to the audio via dynamic time warping (DTW) (Müller, 2015) guided by the transcriber’s predicted note probabilities, and the alignment with the minimum DTW distance is retained as the annotation. The reference MIDI for each excerpt was derived from the corresponding violin MusicXML score. To improve alignment robustness, we applied pitch shifting as data augmentation, generating shifted variants from 5 to +5 semitones for each track. After automatic annotation, the violin note annotations were reviewed by a musically trained annotator and manually refined to improve overall alignment accuracy. During manual refinement, clear onset or offset misalignments were corrected by aligning them to the closest audible note attacks. For trills, we kept the annotations aligned with the notated principal notes indicated in the score rather than marking each rapid alternation as a separate note since the individual note boundaries are difficult to define reliably.

Each violin note annotation consists of onset time, offset time, and MIDI pitch number. Due to the nature of the automatic alignment process, offset estimates are inherently less reliable than onset and pitch values. Velocity values are fixed at 127 throughout the dataset.

3.5 Generating beat annotations

Beat annotation in Western classical ensemble music is particularly challenging due to the absence of percussive instruments and the expressive tempo fluctuations common in the genre. While subtle onset discrepancies between instruments are natural, the ensemble generally adheres to a unified underlying tempo. Due to this complexity, accurate annotation requires contextual musical knowledge and remains a time‑intensive process. Therefore, we adopt a semi‑automated approach by establishing a set of manual annotation guidelines that consider musical context, in addition to extracting automatic alignment based on various resources within the dataset.

We build on the annotation workflow introduced in Foscarin et al. (2020) as a basis, adapting it to the specific characteristics of the ensemble music. Rather than creating separate beat annotations for each instrument, we generate a single global beat annotation shared across both instruments.

We initially extract beat annotations from the piano part using the .match file produced by score‑to‑MIDI alignment, under the assumption that the performers were sufficiently synchronized to yield accurate beat positions. In practice, however, we observed micro‑asynchrony not only in highly expressive sections, such as cadenzas, but also throughout a wide range of passages, which required extensive manual correction to ensure temporal consistency. Moreover, determining the precise tempo location in ensemble music often requires deep musical insight and contextual judgment. To address these challenges, we established the following guidelines:

  • When two performers begin playing together following a visual cue, such as starting the first notes of a piece or resuming after a long pause, the midpoint between the two onsets is set as the beat.

  • When two instruments play the melody and rhythm parts separately, the beat is determined based on the onset of the rhythm part. In most cases, the piano plays the rhythm, but occasionally the violin takes on this role.

  • For violin solo passages where the piano does not establish a clear tempo, the beat is determined based on the onset of the solo instrument.

  • Notes that occur before the beat, such as arpeggios and grace notes, are not considered beats.

  • If one instrument plays incorrectly and slightly disrupts the tempo, the beat is initially determined as the midpoint between the correct beat position of the other instrument and the incorrect note, but is ultimately adjusted based on the musical context.

Additionally, each beat is annotated with its type (downbeat or beat), any time signature changes, and instrument labels of boolean flags that indicate whether the instrument is active; solo passages are explicitly marked, and an instrument is only labeled as resting when it is clearly silent.

4 KRAISLER Dataset Overview

4.1 Dataset structure and components

The KRAISLER dataset consists of multi‑track audio recordings, piano performance MIDI, digital and image formats of musical scores, piano note‑level performance alignment, and annotations. Table 3 summarizes the specific components available for each excerpt, and Figure 2 provides an overview of the structure of the dataset and the relationships among its different modalities.

Table 3

Dataset components for each excerpt.

Data TypesComponentsFormat
Audio (dry/studio/hall)Piano.wav
Violin
Mixture
MIDIPiano.mid
ScoreScore Image.pdf
Score XML.musicxml
MIDI‑score alignmentPiano.match
Note annotationsViolin.csv
Beat annotationsMixture.csv
Figure 2

An example from the KRAISLER dataset with the piece Kreisler’s Liebesfreud, including (a) a piano–violin score excerpt, (b) audio waveforms with beat positions by red lines (solid: downbeat, dash: beat), and (c) a piano roll showing directly captured piano MIDI and estimated violin note annotations.

The dataset comprises 20 piano and violin duets from the Western classical repertoire, featuring works by 13 notable composers. Each excerpt is 1–2 min long, totaling approximately 25 min of audio. The dataset includes stereo recordings of the piano, mono recordings of the violin, and stereo mixtures of the two, all captured at standard CD quality (44.1 kHz, 16‑bit). For each piece, .wav files are provided under three acoustic conditions: dry, studio, and hall.

The musical scores are provided as MusicXML files and corresponding PDF files rendered from them. Each score contains only the performed excerpt and is available in both separate parts for each instrument and a full score. Repeated sections in the original score are fully expanded in the MusicXML files. We include score‑to‑MIDI alignment results in the .match format, which can be loaded via the partitura library7.

We provide note annotations for the violin that include onset and offset times, MIDI note numbers, and velocity information. Additionally, our dataset provides annotations consisting of the list of each beat in seconds; its type as a downbeat or beat; time signature; and instrument labels indicating whether each beat belongs to piano, violin, or both.

4.2 Dataset statistics

4.2.1 Pitch histograms

Figure 3 shows the pitch histograms of piano (upper part) and violin (lower part) note events in our dataset. The x‑axis represents individual pitches labeled using scientific pitch notation, spanning the full 88‑key piano range from A0 to C8. The height of each bar reflects how frequently each pitch was played.

Figure 3

Pitch histograms of piano and violin note events in the KRAISLER dataset. White and black bars correspond to the piano’s white keys and black keys, respectively. Hatched bars in the lower portion correspond to the violin part.

Piano note events are primarily concentrated in the middle register, peaking at D4, with fewer piano note events in the extreme registers below F1 and above C6. In contrast, the violin occupies a relatively higher pitch range than the piano, with the highest event frequency observed at D5. This distribution reflects the typical role structure in violin–piano duets, in which the piano provides harmonic and rhythmic support in the middle and lower registers, while the violin primarily carries the melody in the higher register.

4.2.2 The number of notes per second

Figure 4 shows the piano and violin note density for each piece. We estimated the total note activation time of the piano and violin for each piece using the onset and offset information from the note annotations. For each instrument and piece, the number of notes per second was computed as the total number of notes divided by its corresponding active duration.

Figure 4

Note density in notes per second for the piano and violin of each piece in the KRAISLER dataset.

A higher value of piano notes per second indicates either a faster tempo or a greater degree of polyphony within the performance, whereas a lower value suggests a slower tempo or a more monophonic texture. In contrast, for the violin, the higher number of notes per second generally reflects a faster tempo. However, within a single piece, even when the piano exhibits a high note density, the violin does not necessarily show a similarly high density. For example, in Track 2, the violin primarily carries the main melody with sustained notes, resulting in a low value, whereas the piano exhibits a high value due to its richly polyphonic texture and ornamented accompaniment.

4.2.3 Note activation time

We measured the activation time of the piano and violin for each piece. Although the repertoire was selected to include substantial sections where both instruments perform together, differences in activation time ratios are observed in several tracks. As shown in Figure 5, Tracks 11 and 12 show that the piano plays for nearly the entire duration, whereas the violin appears only in certain passages, resulting in a relatively lower activation ratio for the violin.

Figure 5

Note activation time of each piece in the KRAISLER dataset.

Conversely, there are also pieces in which the violin exhibits a higher activation ratio than the piano. In Track 4, the piano shows a lower activation ratio than the violin because there are passages where the piano rests while the violin plays alone. In this case, the computed activation ratio accurately reflects the actual performance structure. In contrast, in Track 3, the piano activation time ratio is approximately 50%, even though the piano plays almost continuously throughout the piece. This discrepancy arises from the frequent use of staccato articulation, which shortens the computed note durations and results in an underestimation of the actual performance continuity.

It is also important to note that the accuracy of the offset information used to compute activation time differs between instruments. The piano activation time is calculated directly from captured MIDI data. Because the note offsets and pedal information are highly accurate, the resulting activation time estimates are reliable. In contrast, the violin note data were obtained via score–audio alignment, so offset accuracy depends on the alignment algorithm. As the estimated offsets are not perfectly precise, the violin activation time may also include some error.

4.2.4 Tempo distribution

Expressive classical performance is characterized by continuous beat fluctuations throughout a piece. To explore this in our dataset, we first extracted beat timestamps and converted the inter‐beat intervals into instantaneous tempi (beats per minute). Following prior works (Carter and von Appen, 2025; Schreiber et al., 2020), we compute the coefficient of variation (cvar) to measure the relative tempo variability:

cvar=σIBIμIBI

where μIBI and σIBI denote the mean and standard deviation of inter‑beat intervals, respectively. Figure 6 compares tempo curves of two pieces with contrasting cvar values. Kreisler’s Preghiera (cvar=0.29) shows substantial tempo fluctuations, with the pianist employing rubato throughout the passage. In contrast, Beethoven’s Violin Sonata No. 5 (cvar=0.05) maintains a steady pulse, as expected from the Allegro marking. Figure 7 shows boxplots of tempo variability for all pieces in our dataset. The cvar values range from 0.05 to 0.29, indicating that the dataset captures diverse levels of expressive tempo variability.

Figure 6

A comparison of the tempo variability of two pieces in our dataset. (a) 2. Kreisler’s Preghiera (cvar=0.29) (b) 13. Beethoven’s Violin Sonata No. 5 (cvar=0.05).

Figure 7

Tempo distribution for each piece with coefficient of variance (cvar).

5 Applications to MIR Tasks

5.1 Piano transcription on solo recordings

We first applied several state‑of‑the‑art transcription models to our dataset, focusing on piano solo performances. The models were trained on the MAESTRO dataset (Hawthorne et al., 2019) and evaluated on the solo piano tracks of our dataset.

We trained the Onsets and Frames model (Hawthorne et al., 2018), a state‑of‑the‑art piano transcription model with a convolutional neural network–recurrent neural network (RNN) two‑branch architecture that jointly predicts note onsets and frame‑wise pitch activations, using a batch size of 16 and hop length of 512 (32 ms), with a training limit of 100,000 iterations. The learning rate followed the default setting provided in the widely used PyTorch implementation8. The HPPNet model (Wei et al., 2022), a compact model incorporating harmonic‑aware representations via harmonic dilated convolutions and frequency‑grouped RNNs, was adopted using the official implementation9, using a batch size of 8, hop length of 320 (20 ms), and a maximum of 250,000 training iterations. For both models trained on the MAESTRO v3 dataset, the checkpoint with the highest note F1‑score on the validation set was selected for evaluation.

In addition to the models we trained from scratch, we evaluated two state‑of‑the‑art pretrained transcription models. The high‑resolution transcription model proposed by Kong et al. (2021) regresses precise onset and offset times from log‑mel spectrogram inputs by predicting the continuous time distance to the nearest onset or offset for each frame. It is publicly available as a pip‑installable package10 with the released model trained on the MAESTRO v2 dataset. Yan and Duan (2024) introduced a piano transcription model, named Transkun, based on semi‑Markov conditional random fields that directly segments audio into variable‑length note events, using scaled inner product scoring for efficient interval‑level prediction. The model was evaluated in two configurations. The first uses the pretrained model distributed via the pip‑installable package, which, according to the official code11, was trained on MAESTRO v3 with data augmentation. The second is a pretrained variant trained on the same dataset without any data augmentation, allowing for a more direct comparison under training conditions consistent with the other models.

We evaluated the models using standard frame‑based and note‑based metrics provided by mir_eval (Raffel et al., 2014), with an onset tolerance of 50 ms, an offset ratio of 0.2, a velocity tolerance of 0.1, and onset and frame thresholds of 0.5.

Table 4 summarizes the automatic piano transcription results on the piano solo tracks of our dataset. Among the evaluated models, the Transkun model with augmentation achieved the best overall performance across all metrics, with an F1‑score of 90.1% for the frame‑level metric and 98.5% for the note onset metric. It also outperformed others achieving 79.2% when offset was included and 78.3% when both offset and velocity were considered. The Kong model and the Transkun model without augmentation performed similarly across all metrics, with only a slight decrease observed for the Kong model in the note‑with‑offset setting.

Table 4

Automatic piano transcription results evaluated on the piano solo tracks of the KRAISLER and MAESTRO datasets.

ModelMAESTROFrameNote OnsetNote w/ OffsetNote w/ Offset & Vel.
v2v3Aug.PRF1PRF1PRF1PRF1
KRAISLER piano solo tracks
OaF97.766.678.797.386.691.366.859.963.065.158.561.5
HPPNet97.675.584.797.595.896.671.270.070.669.168.068.5
Transkun96.184.589.695.797.396.577.278.477.874.976.075.4
Kong93.186.689.496.896.196.477.276.676.975.775.275.4
Transkun_Aug97.884.290.199.797.398.580.178.379.279.277.578.3
MAESTRO piano solo tracks
Transkun95.895.095.499.597.298.394.692.493.594.191.992.9

Although the HPPNet, Transkun, and Kong models all achieved comparable performance in note onset F1‑scores exceeding 96%, the HPPNet model showed relatively lower performance in frame‑level and note‑with‑offset metrics, yet still consistently outperformed the Onsets and Frames model. These results highlight the value of our dataset as a new benchmark for evaluating the out‑of‑distribution generalization performance of transcription models trained on the MAESTRO dataset.

We compare piano transcription performance on the KRAISLER dataset and also report results on the MAESTRO test split, using the Transkun model trained on the MAESTRO training split. The note onset F1‑scores show only marginal differences between the two datasets, while the frame F1‑score exhibits a performance drop of approximately 6% on our dataset. In contrast, the note‑with‑offset F1‑scores suffer a substantially larger degradation. These results indicate that our dataset effectively serves as an out‑of‑distribution test set and provides a valuable new benchmark for evaluating the robustness of solo piano transcription models.

5.2 Piano transcription on duet recordings

Following the solo transcription experiments, we focused on all the mixtures of the piano–violin duet recordings to evaluate the performance of the aforementioned four transcription models. We used the same models previously applied to the piano solo recordings, all trained solely on the MAESTRO dataset without exposure to any other datasets containing other instruments.

As shown in Table 5, the performance of the models trained with the MAESTRO v3 dataset and without augmentation deteriorated significantly under the duet condition, indicating limited generalization capability to instrument mixtures. More recent model architectures, ranging from Onsets and Frames to HPPNet and Transkun, tend to achieve higher recall by capturing a broader range of note events. However, this improvement often comes at the expense of precision, which tends to decline due to increased false positives. These errors may stem from confusion with violin partials or overlapping harmonic content, ultimately leading to a decrease in the overall F1‑score.

Table 5

Automatic piano transcription results evaluating on all mixtures of piano and violin duets from our dataset.

ModelMAESTROFrameNote OnsetNote w/ OffsetNote w/ Offset & Vel.
v2v3Aug.PRF1PRF1PRF1PRF1
OaF85.344.057.371.174.071.628.429.328.526.727.526.8
HPPNet78.757.866.251.088.463.826.946.633.625.043.431.3
Transkun79.669.573.943.691.358.324.151.532.421.345.528.7
Kong82.485.483.589.691.490.365.366.866.061.262.761.9
Transkun_Aug95.578.685.695.794.595.072.271.471.770.769.970.2

In contrast, the pretrained Kong’s model (Kong et al., 2021) and the Transkun model (Yan and Duan, 2024) with augmentation exhibited relatively strong performance despite some degradation. Specifically, they achieved note F1‑scores of 90.3% and 95.0%, respectively, demonstrating their robustness even in the presence of violin accompaniment.

A possible explanation for Kong’s model robustness lies in its use of the MAESTRO v2 dataset. Notably, the dataset includes six misclassified recordings featuring ensemble performances of piano and string instruments12. These recordings may have inadvertently served as a form of data augmentation in the context of piano‑violin duets.

The Transkun model (Yan and Duan, 2024), on the other hand, was explicitly trained with data‑augmentation techniques provided in the pip‑installable package. These include pitch shifting within ±20 cents and the addition of background noise from the ESC dataset (Piczak, 2015), a corpus of environmental sound recordings, which likely contributed to its superior generalization in duet scenarios.

The impact of data augmentation is particularly evident when comparing the two semi‑CRF configurations. With augmentation, the model improved its note onset F1‑score from 58.3% to 95.0%. Notably, precision more than doubled across all note‑level metrics, indicating that data augmentation significantly reduced false positives while yielding moderate improvements in recall.

5.3 Violin transcription on solo recordings

Extending our analysis beyond piano transcription, we evaluate violin transcription on solo recordings using two violin transcription models, MUSC (Tamer et al., 2023) and VioPTT (Wang et al., 2026). MUSC is a widely adopted violin‑specific transcription model13 based on a Multi‑Stream Conformer architecture, which processes raw audio waveforms to estimate onsets, offsets, semitone‑level pitch frames, and high‑resolution f0 representations. The model was trained on 34 h of solo violin performances, including the Violin Etudes dataset (Tamer et al., 2022), using a weakly‑supervised learning strategy based on iterative audio‑score alignment.

VioPTT (Wang et al., 2026) is a violin‑specific transcription model14 that extends note‑level violin transcription by jointly estimating note pitch and timing information with violin playing‑technique labels from audio. It trains its transcription module on the MOSA dataset (Huang et al., 2024) and constructs the MOSA‑VPT dataset to provide synthetic supervision for playing‑technique prediction.

We evaluated the solo violin tracks in our dataset using the released pretrained MUSC and VioPTT checkpoints without additional fine‑tuning. We then compared these results with the performance on publicly available datasets containing violin recordings, as available for each model: URMP (Li et al., 2018), Bach10 (Duan and Pardo, 2011), and MOSA (Huang et al., 2024). Unlike URMP and Bach10, which are ensemble datasets containing violin parts, MOSA consists of separate solo piano and solo violin performances with aligned audio and note‑level semantic annotations. Since MOSA was used during VioPTT training, we exclude it when evaluating with VioPTT.

For both models, we compute standard note‑level transcription metrics using mir_eval (Raffel et al., 2014), including Precision (P), Recall (R), and and F1‑score (F1) with a pitch tolerance of 50 cents, an onset tolerance of 50 ms, and an offset ratio of 0.2, as well as F1‑score without offset (F1no).

As shown in Table 6, MUSC achieves the best performance on the URMP dataset, followed by Bach10 and MOSA with progressively lower scores. In contrast, its performance on our dataset is the lowest, with an F1 of 40.9% and an F1no of 60.5%. VioPTT shows the same overall trend but achieves higher performance on KRAISLER, with an F1 of 56.7% and an F1no of 75.6%. This suggests that VioPTT’s violin‑specific and technique‑aware training objective, together with synthetic data augmentation, improves robustness on expressive solo violin recordings, although a clear gap remains relative to URMP and Bach10.

Table 6

Violin transcription results using the MUSC (Tamer et al., 2023) and VioPTT (Wang et al., 2026) transcription models on solo violin tracks from KRAISLER and representative datasets containing violin recordings.

TestsetMUSCVioPTT
PRF1F1noPRF1F1no
URMP86.583.184.693.086.183.684.593.1
Bach1065.064.864.877.068.171.869.979.5
MOSA59.457.658.372.2
KRAISLER39.443.740.960.560.553.856.775.6

Since all datasets consist of either clean single‑instrument recordings or bleed‑free multi‑instrument recordings, the observed performance differences are more likely attributable to variations in the technical demands of the repertoire and note‑boundary clarity rather than to recording conditions. The URMP and Bach10 datasets feature relatively straightforward melodic lines with rhythmically regular phrasing, where note onsets and offsets tend to be acoustically well‑defined. The MOSA dataset exhibits greater musical complexity, leading to a moderate performance drop.

The lower performance on our dataset can be attributed to the expressive nature of the repertoire, which includes concerto and sonata movements characterized by substantial rubato, dynamic shaping, and idiomatic bowing techniques such as prolonged sustain and smooth legato transitions that blur acoustic note boundaries. The reduction in F1 is particularly pronounced compared with other datasets. This likely stems not only from the expressive characteristics of the repertoire but also from the limited reliability of the offset annotations. In the violin note‑annotation process, note onsets and offsets are first automatically estimated from model predictions and then manually refined by a musically trained annotator. Since expressive playing often obscures note endings, offset estimation is inherently difficult, which may introduce uncertainty into the offset annotations.

5.4 Music source separation

We also explored the application of our dataset in piano–violin source separation. Currently, notable source separation methods (Choi et al., 2021; Défossez et al., 2019; Hennequin et al., 2020; Jansson et al., 2017; Stöter et al., 2019) primarily focus on isolating four standard stems—vocals, drums, bass, and other accompaniment, with most studies relying on the MUSDB18 dataset (Rafii et al., 2017). However, in this framework, piano and violin are both categorized under others, making it impossible to separate them as distinct sources.

Some notable studies extend source separation to include the piano instrument beyond the standard four‑stem separation. Manilow et al. (2020) approach music source separation by separating into piano, bass, drums, guitar, and strings based on the Slakh dataset (Manilow et al., 2019), the MAPS dataset (Emiya et al., 2010), and the GuitarSet (Xi et al., 2018). Chiu et al. (2020) tackles violin–piano source separation by applying mixing‑specific data‑augmentation techniques. While effective in design, their study is limited by the lack of clean duet data for evaluation, relying on only six songs from the MedleyDB (Bittner et al., 2014) dataset where piano and violin are played simultaneously without severe stem leakage. Özer and Müller (2024) explores piano source separation in classical orchestral music using the Piano Concerto Dataset (Özer et al., 2023), making it highly relevant to our study on piano–violin source separation.

In this study, we applied the HDemucs model proposed by Özer and Müller (2024) to our dataset to evaluate its effectiveness in piano–violin source separation. The pretrained model, released on the authors’ official GitHub repository15, was trained on piano concerto recordings and supports piano source separation. Since our study focuses on piano–violin duets, the violin signal is estimated as the residual after subtracting the separated piano from the mixture. Separation quality was evaluated using three standard metrics: signal‑to‑distortion ratio (SDR), signal‑to‑interference ratio (SIR), and signal‑to‑artifacts ratio (SAR), as implemented in the mir_eval library. The metrics were computed over the entire duration of each track, and the final scores were obtained by averaging the results across all tracks.

Table 7 compares the source separation performance for piano and accompanying signals on our dataset and the Piano Concerto Dataset (PCD) (Özer et al., 2023). On our dataset, the piano and violin tracks achieved SDRs of 9.48 and 9.31 dB, SIRs of 21.78 and 14.90 dB, and SARs of 9.87 and 11.09 dB, respectively. In contrast, the PCD reports SDRs of approximately 8.52 dB for piano and 4.92 dB for orchestral accompaniment, SIRs of 12.31 and 10.69 dB, and SARs of 11.60 and 7.02 dB, with notably lower performance for the orchestral stem. This can be attributed to the relative difficulty of separation, as mixtures with full orchestral accompaniment are inherently more complex than those involving only a single violin.

Table 7

Comparison of piano source separation results in classical music.

DatasetInst.SDRSIRSAR
KRAISLER (Ours)Piano9.48 ± 2.6421.78 ± 4.179.87 ± 2.59
Other (Violin)9.31 ± 2.6014.90 ± 4.0211.09 ± 1.98
PCD (Özer et al., 2023)Piano8.52 ± 3.9712.31 ± 3.6611.60 ± 4.46
Other (Orchestra)4.92 ± 2.9010.69 ± 3.547.02 ± 2.73

5.5 Audio‑to‑score alignment

Audio‑to‑score alignment is an MIR task that aligns music recordings with corresponding scores, and it typically falls into two categories: offline approaches, where the full performance is known, and online (or causal) score‑following approaches that incrementally align in real time. To obtain a common feature representation between score and performance, we first render each MusicXML score as audio via FluidSynth16, using a piano soundfont for both approaches. We then extract constant‑Q transform (CQT) features for the offline method and chroma features for the online method from both the synthesized and recorded audio, each chosen for best alignment performance. For the offline method, we use the DTW provided by librosa (McFee et al., 2015). For the online method, we use an online DTW–based score‑following algorithm proposed by Arzt and Widmer (2010), as implemented in pymatchmaker (Park et al., 2025).

Table 8 shows the beat‑level alignment results using both methods. Alignment was evaluated using average absolute error (AAE) with standard deviation (σ), median absolute error (MAE), and alignment rate within temporal thresholds of 0.3, 0.5, and 1.0 s, as implemented in pymatchmaker. The offline method remains robust across all acoustic conditions, with only marginal differences among the three versions. In contrast, the online method shows clear progressive degradation from the dry through the hall version. In the dry and studio conditions, the online method yields comparable results, with MAE around 139–154 ms and approximately 78% coverage within a 0.5‑s window. With hall reverb, performance degrades to an MAE of approximately 171 ms and coverage below 76%. While the offline method is also slightly affected by reverberation, with MAE increasing from 27 to 30 ms, the online method is considerably more sensitive, as reverberation further obscures the already soft onsets of the violin relative to the percussive piano attacks.

Table 8

Beat‑level audio‑to‑score alignment results for piano– violin ensemble mixes across dry, studio, and hall reverb conditions.

MethodAAE (ms) ±σMAE (ms)Alignment Rate (%)
≤0.3 s≤0.5 s≤1.0 s
[Dry mix]
Offline55.4 ± 108.627.296.698.299.7
Online294.0 ± 378.9138.968.277.988.8
[Studio mix]
Offline55.0 ± 105.129.596.998.499.6
Online304.7 ± 381.7154.367.577.488.9
[Hall mix]
Offline61.4 ± 117.329.996.698.299.5
Online318.8 ± 376.7170.563.875.387.8

5.6 Beat tracking

While most existing beat‑tracking models have been primarily developed and evaluated on percussive genres, the Western classical duets in our dataset, characterized by wide expressive tempo deviations, present a relatively challenging scenario. Chiu et al. (2022, 2023) proposed a set of alternative post‑processing trackers and a framework for evaluating metric‑level switching in expressive music, which they applied to the Maz‑5 (Grosche et al., 2010) and ASAP (Foscarin et al., 2020) datasets.

We benchmarked several representative beat tracking models: a temporal convolutional network (TCN) (Davies and Böck, 2019), which uses dilated convolutional layers; All‑in‑One (Kim and Nam, 2023), a joint metrical– functional structure analysis model based on a transformer with neighborhood attention; and a classic RNN‑based beat tracker (Böck et al., 2016) implemented in the madmom library17. We further applied a set of post‑processing trackers (PPTs) introduced in Chiu et al. (2023), including the dynamic Bayesian network (DBN), simple peak picking (SPPK), and predominant local pulse‑based dynamic programming (PLPDP). In our experiments, we used both studio and hall versions and report results averaged across these two conditions.

As shown in Table 9, the TCN and All‑in‑One models, though representative models evaluated on existing datasets, do not perform well in our dataset. Instead, most of the improvements come from the use of PPTs combined with the RNN backbone, where the peak picking strategy (SPPK) yields the highest overall F1‑score, while PLPDP achieves the best recall. This pattern underscores that pre‑trained architectures alone are insufficient to handle the expressive tempo variability of classical ensemble music; rather, most of the performance gains are attributed to the post‑processing stage. These findings highlight the need for beat tracking models that explicitly account for the characteristics of expressive classical duets, such as the absence of percussion and the presence of highly variable tempo, which pose an ongoing challenge for future research.

Table 9

Beat‑tracking results of the KRAISLER dataset.

ModelAccuracy (%)
PRF1
TCN (with DBN)21.243.127.2
All‑in‑One (with DBN)67.955.758.3
RNN+DBN68.570.167.1
RNN+PLPDP53.192.665.3
RNN+SPPK70.776.171.0

6 Conclusions

We presented KRAISLER, a multi‑track dataset of piano and violin duet recordings under realistic ensemble conditions. It features 20 curated classical pieces, focusing on segments where piano and violin are played together, which were recorded in acoustically isolated rooms as high‑fidelity audio. Each piece includes synchronized audio stems with multiple acoustic conditions, including the original dry recordings and two artificial acoustics: studio and concert hall reverberation. In addition, the dataset includes directly captured MIDI data with the Disklavier piano, aligned musical scores, note annotations, and beat annotations.

We demonstrate the applicability of the dataset by evaluating it on various MIR tasks, including piano transcription in solo and duet scenarios, violin transcription, music source separation, score alignment, and beat tracking. Its synchronized multi‑track format enables systematic benchmarking of MIR models in expressive and realistic ensemble conditions, while certain tasks reveal the challenges posed by reverberant acoustics and the non‑percussive nature of expressive classical performances.

To the best of our knowledge, this is the first publicly available multi‑track dataset based on real expressive performances, containing separate piano and violin stems, high‑precision MIDI data captured from a Disklavier piano, as well as score alignment, note annotations, and beat annotations. By releasing this dataset, we hope to provide a useful basis for studying and evaluating MIR models in ensemble contexts that reflect real‑world musical settings; establish a new benchmark for multi‑instrument transcription tasks; and support a broader range of research areas, including tempo modeling for ensemble music, automatic accompaniment, and acoustic tasks such as dereverberation.

7 Ethical Considerations

This study involved audio recordings of musical performances provided with informed consent from the performers. The recordings do not contain any personally identifiable information. Therefore, ethical approval was not required according to institutional guidelines.

Notes

[1] KAIST stands for the Korea Advanced Institute of Science and Technology, a science and technology university in South Korea.

Acknowledgments

We thank Seyoung Kim (violin) and Soobin Kim (piano) for recording the performances in our dataset. We also thank Wonil Kim and Simon Schwär for repeatedly providing the authors with professional guidance on audio mixing.

Data Accessibility

The KRAISLER dataset is available at https://zenodo.org/records/21082251.

Funding Information

This work was supported by the Culture, Sports, and Tourism R&D Program through the Korea Creative Content Agency grant funded by the Ministry of Culture, Sports and Tourism in 2025 and 2026 (Development of Core Technologies for Copyright Verification of AI‑Generated and Deepfake Music, RS‑2025‑02216483; Development of AI‑based Image Expansion and Service Technology for High‑Resolution (8 K/16 K) Service of Performance Contents, RS‑202400395886; Contribution Rate: 25% each); by the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT) (RS‑2023NR077289); and by Institute for Information & Communications Technology Planning & Evaluation (IITP) grant funded by the Korea government (MSIT) (RS‑2019‑II190075, Artificial Intelligence Graduate School Program (KAIST)).

Competing Interests

JN is a member of the Editorial Board of this journal and had no involvement in the review process or editorial decision‑making. All other authors declare no competing interests.

Authors’ Contributions

HK and JP contributed equally to this work and share first authorship. HK conducted the piano transcription and source separation experiments, while JP carried out score alignment and beat tracking. Both led the construction of the dataset and the writing of the manuscript. SL conducted the violin transcription experiments, led the writing of the corresponding section, and contributed to dataset construction. TK supported the piano transcription experiments. SW contributed to dataset creation. JN supervised the work, provided feedback throughout the project, and contributed to editing the manuscript. All authors discussed the results and approved the final manuscript.

DOI: https://doi.org/10.5334/tismir.338 | Journal eISSN: 2514-3298
Language: English
Page range: 456 - 473
Submitted on: Sep 1, 2025
Accepted on: May 18, 2026
Published on: Aug 6, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Hyemi Kim, Jiyun Park, Sein Lee, Taegyun Kwon, Sunjae Won, Juhan Nam, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.