Skip to main content
Have a personal or library account? Click to login
Distorted Realities: Classifying Extreme Vocals Between Harmony and Noise – A Machine–Human Evaluation of Vocal Confusion Patterns Cover

Distorted Realities: Classifying Extreme Vocals Between Harmony and Noise – A Machine–Human Evaluation of Vocal Confusion Patterns

Open Access
|Jul 2026

Abstract

Extreme vocal techniques like growling and shrieking are common in metal and related genres, giving the music a characteristic intensity. Although widely used, these techniques have received limited attention in Music Information Retrieval (MIR) and musicology, particularly in fine‑grained categorisation. Their niche status and marginalisation have led to limited representation in digital music libraries. Since current classification systems prioritise dominant subgenres due to a lack of detailed vocal categorisations, extreme vocal styles are often oversimplified or misrepresented, hindering tagging, retrieval, and analysis. Automation could enable consistent and nuanced vocal annotation, supporting the technical goals of MIR and analytical aims of musicology. Foregrounding vocal technique as a classification axis contributes to a richer understanding of vocal performance and genre formation. To assess machine learning as a reliable alternative to human judgment, this study developed a coarse frequency‑based qualitative taxonomy from interviews with 11 extreme vocal practitioners and compared human and machine performance using the Extreme Metal Vocals Dataset (EMVD), including a listening test with 158 expert participants. Perceptual results achieved 76.2% Unweighted Average Recall (UAR), while a Support Vector Machine (SVM) model trained on ComParE features achieved a UAR of 90% on average, using singer‑independent 3‑fold group cross‑validation. Feature group, ablation, and human prediction analyses suggest that the human‑machine gap reflects differences in cue prioritisation and classification performance. Results illustrate the feasibility of machine learning to annotate vocal styles, while showing that human confusion patterns remain informative for category boundaries. This highlights how MIR methods could facilitate analysis of underrepresented genres.

DOI: https://doi.org/10.5334/tismir.310 | Journal eISSN: 2514-3298
Language: English
Page range: 309 - 328
Submitted on: Jun 30, 2025
Accepted on: May 13, 2026
Published on: Jul 16, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Xuhong Qiu, Emilia Parada‑Cabaleiro, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.