Skip to main content
Have a personal or library account? Click to login
Distorted Realities: Classifying Extreme Vocals Between Harmony and Noise – A Machine–Human Evaluation of Vocal Confusion Patterns Cover

Distorted Realities: Classifying Extreme Vocals Between Harmony and Noise – A Machine–Human Evaluation of Vocal Confusion Patterns

Open Access
|Jul 2026

Figures & Tables

Figure 1

Log‑mel spectrograms of the four investigated vocal techniques (Clear Voice, Black Shriek, Death Growl and Hardcore Scream) performed by Singer ID 10 (male) for ‘Mid_a’ (vowel [a] in a mid‑frequency range). Each panel shows the first 4 s of the recording. The x‑axis shows time (s) and the y‑axis shows the mel frequency with tick labels in Hz. The white box indicates the main band (20th–90th percentile energy interval).

Figure 2

Group‑level estimated marginal means for max_ frequency. Error bars indicate 95% confidence intervals.

Table 1

Distribution of samples across categories (i.e. vocal techniques), considering both the EMVD and the selected samples used for the perceptual versus ML experiments. The number of samples and durations in seconds are given.

CategoryEMVDSelected
SamplesDuration (s)SamplesDuration (s)
Clear Voice3032367.811186.7
Black Shriek1911360.720321.9
Death Growl2221551.221356.5
Hardcore Scream2301677.919266.6
Total9466957.6711155.1
Figure 3

Listening‑test interface used in the human‑perception experiment. Genre labels were removed to minimise genre‑related bias.

Figure 4

Relationship to extreme vocal techniques.

Figure 5

Familiarity with extreme vocal techniques.

Figure 6

Genre preferences related to extreme vocals.

Figure 7

Sample distribution across the three independent experiments. Each experiment presents a different test set (together, totalling the 71 samples used in the perceptual study). For training and evaluation sets, μ±σ is given across the singer‑independent three‑fold group cross‑validation (by singer ID) performed for optimisation.

Table 2

Distribution of the 71 samples included in the test set across vocal techniques. To guarantee a singer‑independent ML setup, the samples were distributed across three independent experiments with unique singers in the test set.

TechniqueExperiment 1Experiment 2Experiment 3
Clear Voice443
Black Shriek668
Death Growl597
Hardcore Scream937
Total242225
Figure 8

Human perception. Confusion matrix for the identification of the four extreme vocalisations.

Figure 9

ML. Confusion matrix for the classification of four extreme vocalisations.

Figure 10

Per‑class broad ComParE feature group importance for the EMVD label model (mean ± standard deviation across folds), with distribution shown within class ranks.

Figure 11

Per‑class broad ComParE feature group importance for the human‑prediction model based on perception labels (mean ± standard deviation across folds), with distribution shown within class ranks.

Figure 12

Comparison of broad ComParE feature group importance between the baseline classification model and the human‑prediction model, reported as mean ± standard deviation across folds with within‑model ranks.

Figure 13

Ablation analysis: restricted model (based on Jitter, Shimmer and HNR). Confusion matrix.

Figure 14

Ablation analysis: complementary model (ComParE excluding Jitter, Shimmer and HNR).

DOI: https://doi.org/10.5334/tismir.310 | Journal eISSN: 2514-3298
Language: English
Page range: 309 - 328
Submitted on: Jun 30, 2025
Accepted on: May 13, 2026
Published on: Jul 16, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Xuhong Qiu, Emilia Parada‑Cabaleiro, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.