1. Introduction
Extreme vocal techniques are widely employed in certain musical genres, including hardcore, punk, metal and various traditional styles, to enhance musical expressiveness and evoke strong emotional responses. Often learned by imitation (Erbe, 2014) and informally distinguished through practice‑based labels used by vocalists (see Appendix 2 for all vocalist interviews), systematic classification could aid understanding and support multiple stakeholders. These include preventing vocal damage, aiding therapy, improving tagging systems and supporting musicological studies. In particular, diagnostic ambiguities arise from the acoustic overlap between artistic extreme vocal techniques and pathological voices.
Currently, extreme vocal–classification frameworks are subjective and historically developed. This has resulted in multiple taxonomies being used by researchers (for discussion, see Tailleur et al., 2024). These taxonomies often lack consistency and reproducibility. They are also insufficiently grounded in measurement, which makes category boundaries difficult to define and compare across perceptual, clinical and computational contexts. Moreover, many practitioner labels used in scene discourse (e.g. ‘Fry Scream’, ‘Tunnel Growl’ and ‘Black Shriek’) have clear meaning within a given community, but are not consistently using measurable acoustic or physiological criteria. This variability can lead to the inconsistent use of labels across studies and can make the resulting categories harder to reuse in clinical and computational contexts.
Research comparing human perception with machine learning (ML) classification shows that artificial intelligence (AI) can effectively recognise patterns associated with perceptual cues, which provides valuable insights into vocal characteristics and can help illuminate perceptual ambiguities in existing classification systems (Stadler et al., 2023). Human versus machine comparisons can illuminate perceptual ambiguities; highlight limitations in existing classification methods; and contribute to the development of more accurate, perceptually aligned models.
A crucial conceptual distinction in this study concerns the difference between perceptual and computational approaches to vocal categorisation. In this context, the perceptual–computational distinction is used as an operational contrast between two categorisation procedures applied to the same audio material: (1) a human forced‑choice listening task and (2) a supervised ML classification based on extracted acoustic features. Moreover, because ML classification is grounded in measurable acoustic features, its outcomes can serve as objective evidence for evaluating vocal taxonomies that are still not fully consolidated.
To this end, we adopt a perceptual versus ML approach to identify four extreme vocal types (Black Shriek, Death Growl, Hardcore Scream and Clean Voice1) in the Extreme Metal Vocals Dataset (EMVD) presented by Tailleur et al. (2024). The research comprises two phases: (1) an in‑the‑wild qualitative evaluation of extreme vocal techniques and (2) an empirical comparative assessment of human perception versus ML performance. In the first phase, in order to understand to what extent current formalised taxonomies align with real‑world practices, we explore how vocalists intuitively categorise extreme vocal techniques based on their practical experience through open‑ended interviews with seven participants. Since this exploratory categorisation lacks the consistency required by ML methods, in the second phase, we systematically evaluate the predefined categories from the EMVD (Tailleur et al., 2024). The listening test was conducted through a forced‑choice task involving 158 expert participants. The ML experiments (using the same audio excerpts assessed by human participants as the test set in order to enable direct comparison) were based on a support vector machine (SVM) fed with ComParE (Eyben et al., 2010) features. By interpreting confusion errors of both human listeners and ML classifiers, we aim to uncover perceptual ambiguities and identify directions for optimising future ML methods. More specifically, while previous research has acknowledged the need for systematic classification, there is still limited evidence on how practitioner terminology aligns with formal dataset labels in extreme vocals. Direct comparisons of human and ML confusion patterns on the same audio material also remain limited, and recurring confusion patterns are not often discussed in terms of category stability, overlap or possible transitional forms. This also provides empirical data to support research on marginal music cultures and reveals how extreme vocal terms are shaped within specific communities, as label meanings and category boundaries are negotiated and stabilised in scene discourse among vocalists and listeners. To facilitate future work and reproducibility, the code and related resources are made available, and the interview materials are included in the appendix.2
2. Related Work
Specialised vocal techniques such as Shriek, Scream, Growl or Pig Squeal arise from non‑traditional vibratory behaviours within the vocal folds and other laryngeal structures (Erbe, 2014; Nieto, 2008; Smialek et al., 2012; Tailleur et al., 2024; Wallmark, 2018). The complexity of these techniques contributes significantly to diversity in timbre and expressiveness, making them widely prevalent in metal, hardcore, punk and related music styles. For instance, vocalists might employ specific techniques to mimic the characteristics of large beasts (Smialek et al., 2012) or produce screams through inhalation (Tailleur et al., 2024).
Different types of extreme vocals rely on distinct vibratory mechanisms involving various laryngeal structures, such as ventricular folds (false vocal folds), vocal folds, aryepiglottic folds and resonating cavities (Erbe, 2014; Eckers et al., 2009; Sakakibara et al., 2004; Smialek et al., 2012). However, the physiological mechanisms underlying these vocalisations vary considerably, which can contribute to similarities with pathological voice characteristics. For example, Drone and Kargyraa in throat singing may involve the synchronous vibrations of ventricular and vocal folds (vocal‑ventricular mode [VVM]), a pattern also observed in certain voice disorders, whereas Vocal Fry primarily involves low‑frequency vocal fold vibrations, often associated with hypofunctional phonation.
Previous work has suggested that auditory and acoustic similarities between Growl, Drone, Vocal Fry and certain pathological voices exist (Sakakibara et al., 2004), cautioning that unclear classification could lead to medical misdiagnosis. Therefore, a well‑structured classification system may help medical practitioners and researchers distinguish between healthy vocal techniques and practices associated with increased vocal health risks. This would not only guide voice rehabilitation but also assist singers in mastering techniques safely, avoiding vocal fold damage. Since different musical genres employ various extreme vocal techniques (Smialek et al., 2012), establishing systematic categories has been shown to aid streaming platforms in recommendation tagging (Nieto, 2013; Tailleur et al., 2024), to assist musicologists in analysing genre‑specific stylistic characteristics and to support advancements in automatic speech recognition (Gentilucci et al., 2018; Ishi et al., 2007; Tailleur et al., 2024) and AI‑based music generation.
The physiological mechanisms and laryngeal behaviours underlying extreme vocal techniques have been examined through various analytical methods, including spectrogram analysis with tools such as AudioSculpt (Smialek et al., 2012), aerodynamic measurements (Guzman et al., 2018), endoscopic imaging, radiographic examinations and acoustic analyses (Caffier et al., 2018; Eckers et al., 2009; Fuks et al., 1998; Kato and Ito, 2013; Lindestad et al., 2001; Titze, 2008). However, while these studies apply advanced technological tools for data collection and analysis, their findings have seldom been integrated into computational models for systematic classification or practical training purposes. This limits direct computational comparison across vocal categories.
Existing research indicates that technological applications, such as audio analysis and real‑time computational feedback, can significantly support the learning and training of singing techniques (Lã, 2012; McCoy, 2014), suggesting that there is potential for similar approaches to be applied to the analysis of extreme vocal techniques. In related sound‑detection research, recent studies have focused primarily on identifying various types of screams and growls (Ishi et al., 2007; Nandwana et al., 2015; Pohjalainen et al., 2011), providing a foundation for future work in automated recognition and classification of extreme vocal styles.
Groundwork formalising extreme vocal classifications includes the Metal Vocal Dataset (MVD) introduced by Kalbag and Lerch (2022), which encompasses 57 songs and focuses on Nieto’s Fry Scream typology (Nieto, 2008, 2013). Building on this foundation, the EMVD dataset by Tailleur et al. (2024) offers a larger and more diverse collection of annotated samples, supporting advances in data‑driven classification approaches. Nevertheless, besides these classification frameworks (Nieto, 2008, 2013; Sadolin, 2012; Tailleur et al., 2024), systematic and data‑driven classification models remain limited, with inconsistent operational definitions across studies and few direct comparisons of human and model confusion on the same audio material.
Despite the work outlined above, methods for automatic recognition of extreme vocalisations remain limited, which prevents the evaluation of shared confusion patterns (Parada‑Cabaleiro et al., 2023). Previous research has suggested that ML can use multiple perceptual cues (auditory and visual), mirroring human perception and thus potentially supplementing pedagogical tools and encouraging self‑assessment in vocal students (Stadler et al., 2023). Comparing ML models with human perceptual performance could illuminate the strengths and limitations of current classification systems, uncover cognitive mechanisms in auditory processing and provide perceptual benchmarks for optimising classification models.
3. Extreme Vocal Categorisation
3.1 Existing taxonomies and classification challenges
Currently, most classification systems in the extreme metal field are based on qualitative, practice‑based observations rather than being grounded in systematically quantified acoustic parameters (Tailleur et al., 2024).3 Sadolin (2012: 177–202) classified some lighter vocal techniques found in extreme metal, such as Distortion, Rattle, Growl and Grunt. Nieto (2008, 2013) introduced the concept of ‘Vocal Effects’ or ‘Extreme Vocal Effects’, categorising extreme vocals into three groups: (1) Growl, common in death metal and other extreme music genres, is highly noisy, with fundamental frequencies rarely perceived, and it involves significant spectral variation; (2) Fry Scream, induced by irregularly spaced glottal pulses during inhalation or exhalation, is commonly known as Pig Squeal and is similar to Growl but with a brighter spectrum; and (3) Roughness, a category that includes techniques that add subtle variations to the vocal tract, produces a harmonically richer spectrum and is more commonly found in hard rock.
Wallmark (2018) likewise summarises distorted vocals in death metal under the umbrella term ‘death vocals’, which includes growls, grunts, pig squeals, screams and deep gutturals. He describes these as predominantly low‑register singing styles characterised by a thick, raspy and heavily distorted vocal quality. He also notes that singers conceptualise their tonal options in broad register categories – namely highs, mids and lows (see Wallmark, 2018: 72). In the MVD (Kalbag and Lerch, 2022), the authors proposed that most screams in modern metal are variations of Fry Scream. Their dataset mainly focuses on the classification of Fry Scream, dividing it into three perceptual categories: high, mid and low. In scene‑informed discourse, Kennedy (2018) divides distorted vocals into three broad categories, growls, screams and roars, and relates these categories to coarse frequency tendencies and genre associations. The author also notes that boundaries between growls and roars can be malleable and often indistinct, with roaring occupying an intermediate region between low growls and high screams or shrieks (Kennedy, 2018).
Tailleur et al. (2024) introduced a more comprehensive classification that not only covers vocal techniques but also emphasises their association with specific extreme subgenres. Their classification is divided into: (1) Vocal Techniques, defined by the ability to replace a clean voice, articulate understandable lyrics and adjust pitch within a musical range (e.g. Black Shriek, Death Growl, Hardcore Scream and Grind Inhale), and (2) Vocal Effects, defined by the incomprehensibility of lyrics and their typical use as embellishments in performances (e.g. Pig Squeal, Deep Gutturals and Tunnel Throat). In this classification, vocal techniques are primarily methods of sound production that can replace clean vocals while maintaining pitch control and lyric articulation. In contrast, vocal effects are embellishments that often obscure lyrics while serving as textural or expressive devices.
In addition to classification methods specific to extreme metal music and related genres, there are other broader classification frameworks for extreme vocals, which may be applicable to this research. For example, classification can be based on anatomical structures and phonatory mechanisms, such as the vibratory phonation of the ventricular folds, vocal folds and aryepiglottic folds mentioned earlier, as well as the taxonomy of Supraglottic Structure Vibrations; see Aaen et al. (2025) for discussion. A classification based on resonating cavities and airflow physiology can also be considered: Extreme vocals can be produced using either exhalation or inhalation, resulting in two distinct styles (Kalbag and Lerch, 2022; Smialek et al., 2012).
Finally, during vocal performance, singers can also modify their phonation either by lowering or raising the jaw, which affects the perceived pitch and timbre of extreme vocals (Smialek et al., 2012) and can be used as a classification criterion.
3.2 Practice‑Based taxonomies: Interview data
In practical contexts, classification among practitioners is often experience‑based. To capture practice‑grounded perspectives, we conducted exploratory interviews with 11 experienced extreme metal vocalists (see Appendix 2 for transcripts). Interviews were conducted remotely in an asynchronous format via online messaging. Participants responded in writing and, when preferred, with short voice messages, allowing them additional time to reflect on the questions and formulate their answers. The first interview used open questions about possible classification principles and phonation strategies. Based on this first interview, in which the vocalist proposed a high‑, mid‑ and low‑frequency–based grouping, we later asked all vocalists how they viewed this classification. Even when one or two participants stated upfront that they themselves already classify by frequency, we still posed the same question to maintain a consistent interview protocol and enable comparable responses across participants. As noted by Fan, the vocalist of Frozen Moon and one of the managers of Pest Productions, the largest black metal label in Asia, during a Personal Interview (see Appendix 2A):
‘Since there are many vocal techniques, I cannot fully grasp them, and an exhaustive organisation is probably not feasible. Apart from blackened death metal, other metal subgenres do not have a mandatory singing style. However, overall, the classification is mainly determined by frequency differences. I generally divide them into high‑frequency distorted screams, mid‑frequency distorted screams and low‑frequency gutturals’. (See Appendix 2A)
Dyingflames, the owner of Diensysian Records and a musician in Holokastrial and R.N.V., expressed a similar opinion:
‘Black vocals, death vocals and hardcore vocals. But in reality, these classifications are also frequency‑based. For example, Growls and Gutturals in death metal are low‑frequency; black metal vocals are mid‑, high‑frequency; and hardcore vocals typically range from mid‑ to high‑frequency or high‑frequency. We rely on auditory perception and musical style to distinguish and classify them’. (See Appendix 2C)
An anonymous black metal singer agreed with their perspective:
‘It’s a very complex subject and depends on the style of music you sing. Different vocals have different frequencies. You can definitely use frequency, but it’s just different ranges of frequency and Hz. [ . . . ] Think of, for example, the Chinese tones. I mean, you can measure scientifically how many Hz are in a certain range, but when it comes to feeling, you can’t measure that’. (See Appendix 2G)
Another vocalist, Jovi, who has extensive experience in cover performances, proposed a slightly different classification system in a personal interview (see Appendix 2E), suggesting that extreme vocals could be categorised into four groups: Recitation, Clean Singing, Melodic Singing and Distortion. Three metal vocalists (20+ years experience) supported frequency‑based classification but noted that fixed frequency boundaries are hard to define. Additional practitioners expressed a more style‑oriented perspective. GUT from Huangquan Records and the death metal band Mvltifission argued that classification should primarily follow stylistic context, since a simple ‘high, mid, low’ scheme is too general and each style itself can contain high, mid and low vocal approaches (see Appendix 2F). DomAs666 from Naked Whipper agreed with this position and considered a style‑based classification more flexible for both musical and instrumental practice, allowing the voice to produce a wider range of sonic effects (see Appendix 2I). In his view, the traditional ‘high, mid, low’ framework is workable but overly rigid for practical artistic use. However, other interviewees held a different opinion and regarded the ‘high, mid, low’ system as more practical overall, arguing that, when the underlying production principles are similar, further subdivision is not necessary.
As Fan noted (see Appendix 2A), there are many vocal techniques, and a fully exhaustive organisation is difficult to achieve. This difficulty is compounded by the fact that, as Brackett (2016) points out in his discussion of the concept of genre, genre boundaries are themselves porous and blurred, and overlap and intersection constantly occur between closely related subgenres (see Chapter 1 in Brackett, 2016 for discussion). A similar pattern can be observed in extreme music, especially in the late 1980s and early 1990s, when the stylistic practices of many bands did not strictly conform to a single genre framework but instead simultaneously displayed features of multiple extreme music styles. For example, in the works of the Brazilian bands Vulcano and Sarcófago, traits associated with thrash metal, death metal and black metal are often intertwined. Similarly, the boundaries between grindcore, punk and metal are far from clear. In the case of Napalm Death, one can hear elements of grindcore, hardcore punk, death metal and thrash metal across different periods or releases; likewise, the early form of grindcore displayed by Extreme Noise Terror also reflects, to a considerable extent, the continuity and hybridity between hardcore punk and metal.
In such contexts of genre overlap or hybridity, the vocal styles used by bands may also vary considerably, depending to a large extent on the particular vocalist’s phonatory habits and technical approach; at the same time, when the other members and the overall compositional orientation remain relatively unchanged, the musical style of the band may not necessarily undergo a fundamental shift as a result.4
This suggests that there is no simple one‑to‑one correspondence between genre labels and vocal labels; a single genre may accommodate markedly different vocal realisations, while similar vocal sound results may also occur in different stylistic contexts. Although GUT and DomAs666 suggested in our interviews that classification could follow stylistic context, this is not always sufficient. Kennedy (2018) notes that genre framing can itself rely on contrasts between clean singing and distorted vocals – for example, in progressive death metal and related contexts. For this reason, we adopt the high‑, mid‑ and low‑frequency–based grouping as a first step.
Based on insights from these interviews and the recommendations of the vocalists, our research group5 analysed 117 audio excerpts of punk, metal and hardcore vocals, provided as part of a seminar in the musicology program at the University of Cologne.6 The excerpts have a duration of 6.2 s on average (range 0.6–22.8 s, N = 117). The simple frequency‑based classification (high, mid, low) was operationalised by assigning excerpts to three frequency groups via Statistical Package for the Social Sciences (SPSS) automatic binning of maximum amplitude frequency (see Section 4.1, Footnote 7). Though seemingly intuitive or subjective, the consistent recurrence of the simple frequency‑based classification across interviews indicates perceptual coherence. Rather than treating it as a fixed taxonomy, we use it as an initial sanity check to explore whether such performer‑informed distinctions align with measurable acoustic features. This perspective allows us to contribute to ongoing efforts to bridge the gap between practice‑based knowledge and computational modelling in extreme vocal analysis. Although vocal techniques and effects are often discussed in terms of acoustic features, some practitioners distinguish them primarily by articulatory technique, regardless of the resulting sound.
Interviews revealed that, while listeners use frequency cues to distinguish vocal types, techniques like ‘Black Shriek’ and ‘Hardcore Scream’ differ in vocal production. These differences are genre‑specific and involve distinct physiological mechanisms (see Section 4 ‘Methodology’). While listeners familiar with the respective genres (e.g. black metal or hardcore) often identify these techniques by ear, performers themselves emphasise that they rely on different vocal‑production strategies. Vocalists learn them through embodied experimentation and recognise them within shared stylistic frameworks. Some techniques (e.g. growl and pig squeal) are documented, while others evolve informally as genre‑specific or hybrid forms. Though rarely formalised in scholarly literature, performers consistently acknowledge them as distinct and reproducible within their respective traditions. As such, performer intuition and stylistic convention are indispensable in understanding the functional boundaries of extreme vocal techniques.
The interview‑derived taxonomy offered insight into how performers intuitively group vocal styles. However, because such groupings are inherently dependent on the vocal range of a given singer and lack cross‑performer standardisation, they were considered unsuitable as ground‑truth labels for computational modelling in the later phase of this study. Instead, we treated them as a perceptually grounded probe that informs the acoustic validation and serves as a reference framework for the ML analyses presented in Section 4.
3.3 Singing techniques
As discussed in the previous section, current extreme vocal‑classification frameworks are largely practice‑based and historically developed, which leads to inconsistencies across taxonomies. To empirically evaluate these categories, we conducted a set of perception and ML experiments using the EMVD by Tailleur et al. (2024). The dataset provides operational class labels for supervised training and evaluation. In the following subsections, we report the human‑perception experiment and the subsequent ML analyses (see Section 4), including the main classification model as well as auxiliary analyses, such as feature‑group importance and ablation tests. A description of the EMVD is provided in Section 4.2.
To better understand the data evaluated and the results achieved, we provide a brief introduction to the four vocal techniques used in this study. Figure 1 displays log‑mel spectrograms generated using the Python library Librosa by McFee et al. (2015). Log‑mel spectrograms were chosen because they align better with human auditory perception, i.e. processing frequency distributions similarly to how humans naturally perceive sound (Jothimani and Premalatha, 2022; Li et al., 2024). The spectrograms show audio from the EMVD recorded by singer ID 10 (male) performing ‘Mid_a’, i.e. the vowel [a] performed in a mid‑frequency range, using each of the four vocal techniques. In the EMVD dataset, Black Shriek, Death Growl and Hardcore Scream are categorised as Vocal Techniques. These are alternative vocalisations to Clean Voice, capable of conveying intelligible lyrics while allowing for stylistic variations. Supraglottic constriction plays a significant role in extreme vocalisation techniques (Eckers et al., 2009; Erbe, 2014; Sakakibara et al., 2004), making it a key distinguishing feature between clean voice and other vocal styles. As clearly visible in the spectrograms (see Figure 1), clean voice can readily be distinguished from other techniques, exhibiting distinct and stable harmonic lines, particularly concentrated in the lower frequency range. For a reproducible comparison across techniques, the white box marks the main band, defined as the 20th to 90th percentile frequency interval of the time averaged mel band energy over the displayed excerpt.

Figure 1
Log‑mel spectrograms of the four investigated vocal techniques (Clear Voice, Black Shriek, Death Growl and Hardcore Scream) performed by Singer ID 10 (male) for ‘Mid_a’ (vowel [a] in a mid‑frequency range). Each panel shows the first 4 s of the recording. The x‑axis shows time (s) and the y‑axis shows the mel frequency with tick labels in Hz. The white box indicates the main band (20th–90th percentile energy interval).
Black Shriek demonstrates comparatively stronger energy in the upper frequency region than the other techniques, with a dense and sustained appearance in the spectrogram. From an auditory perspective, Black Shriek (from now on also referred to as Shriek for simplicity) resembles a harsh, high‑frequency noise like vocal quality with relatively limited tonal salience. This type of sound is widely used in black metal, including canonical bands such as Darkthrone and Emperor, and can also be heard in crossover acts such as Violent Magic Orchestra. In the EMVD metadata, it is classified as an exhaled technique and is generally described as higher than Death Growl (Tailleur et al., 2024). In a personal interview (Appendix 2C), Dyingflames indicated that Shriek and Scream share similar vocal‑production techniques: ‘High[‑]pitched extreme vocals are mostly similar in production. Shriek has a more pronounced grainy texture, and during production, just like with Scream, the mouth is opened in a smile‑like shape, and the throat is compressed’.
Death Growl, on the other hand, is perceived as a low‑frequency sound with minimal human vocal characteristics, resembling animalistic growling. Growl is common in death metal and can be heard in bands such as Deicide, Morbid Angel, Pestilence, Grave, Death and Suffocation. Its spectrogram shows energy predominantly concentrated in the low‑ to low‑mid‑frequency range, appearing dense but somewhat blurred and lacking sharp, high‑frequency features. An endoscopic study (Eckers et al., 2009) revealed two primary ways singers produce this technique: (1) vibration of the ventricular folds accompanied by an anteroposterior constriction of the supraglottic tract and (2) vibration of the aryepiglottic folds, which produces a sharper, squealing sound similar to a Pig Squeal. Erbe (2014) shows that some singers performing Growl do not use vocal fold vibration; instead, they maintain their vocal folds in a rigid state and rely on vibrations from other anatomical structures. In some cases, the epiglottis strikes and makes contact with the arytenoid cartilages. Additionally, mucus or saliva may vibrate during performance, contributing to the roughness of the vocal timbre while simultaneously providing a protective effect for the vocal folds. Wallmark further notes that, during phonation, singers should keep the throat well hydrated, avoid excessive throat tension and sing with diaphragmatic support (see Wallmark, 2018: 75). Dyingflames, known for his proficiency in Growl, explained: ‘Growl is different from the other techniques. The mouth forms an “O” shape, and the vocal folds lower – something like that’ (see Appendix 2C).
Hardcore Scream occupies mid‑ to high‑frequency ranges and retains partial human voice characteristics. The spectrogram clearly shows layered frequency bands from low frequency (approximately 128 Hz) to mid‑frequency (approximately 2048 Hz), reflecting distinct vocal resonance layers. Interviewee B described this vocal technique as ‘the vocal cords coming together, with airflow pushed upward, forcing a little sound through tiny gaps’ (see Appendix 2D), which illustrates the characteristics observed in the spectrogram from an audible point of view. A similar vocal quality can be heard in the music of The Blood Brothers and Slant. It is also present in the vocals of Fuming Mouth, where Mark Whelan shifts between Scream, Growl and Clear Voice in his performances. Finally, Dyingflames added, ‘the sound of Hardcore Scream is drier than Black Shriek’, although he noted that the production technique is essentially the same (see Appendix 2C).
However, the term Hardcore Scream remains open to discussion, as it does not necessarily denote a homogeneous category in scene discourse. Kennedy (2018) describes early hardcore vocals as closer to an overdriven shout and discusses practices such as group vocals, illustrating that hardcore vocal delivery is not reducible to a single technique and that genre context shapes how vocal descriptors are used across hardcore and metal‑related styles. In Kennedy’s account, screaming is often associated with grindcore and black metal contexts, yet our material suggests that label usage and acoustic tendencies do not map one to one across corpora and communities. In particular, Figure 1 and several interview comparisons indicate that the present dataset categories capture distinctions that are context‑dependent and may differ from genre level descriptors used elsewhere (see Appendices 2C, 2D, 2F and 2I). Therefore, we retain the EMVD category names as operational labels for the current study and follow the dataset annotation by treating Scream as Hardcore Scream in the experimental sections.
4. Methodology
4.1 Taxonomy validation
As an initial step, we preprocessed and analysed the 117 audio excerpts provided in the University course to evaluate whether the frequency‑related taxonomy introduced in Section 3.1 and derived from practitioner’s interviews corresponded to measurable acoustic differences.7 Each clip was transformed into the frequency domain using a discrete Fourier‑transform implemented through SciPy FFT (fast Fourier transform); see Virtanen et al. (2020). The transform was applied once to the full‑length waveform (i.e. no short‑time framing), yielding a single‑sided magnitude spectrum in Hz.8 To assess the validity of the interview‑derived categorisation, the samples were divided into three groups, i.e. high‑, mid‑ and low‑frequency, based on their maximum amplitude frequency (max_frequency), defined as the frequency location of the maximum spectral magnitude.
We conducted a one‑way analysis of variance (ANOVA) to compare the group means. Because the assumption of homogeneity of variances was violated (Levene’s test (, )), we used Welch’s ANOVA, which revealed a significant group effect (, ) with a large effect size (, 95% CI [0.31, 0.57]). Post‑hoc pairwise comparisons, evaluated using Games–Howell tests, revealed that all pairs were significant (): the high‑frequency group differed from the mid‑frequency group by about 291 Hz and from the low‑frequency group by about 535 Hz; the mid‑frequency group also differed from the low‑frequency group by about 245 Hz. These results support the validity of our taxonomy‑based frequency groups in terms of maximum amplitude frequency. To complement the analysis based on the maximum‑amplitude frequency, we repeated the validation using the spectral centroid (Hz).9 Since the assumption of homogeneity of variance was again violated (Levene’s test: , ), Welch’s ANOVA was applied. The ANOVA results (Welch’s , , , 95% CI ) indicate a very strong group effect. Games–Howell post‑hoc tests further revealed significant pairwise differences between all groups (all ): the high group exceeded the mid group by approximately Hz and the low group by approximately Hz; the mid group exceeded the low group by approximately Hz. These results closely mirror the max_frequency analysis, which supports the validity of our frequency‑based taxonomy.
The bar chart in Figure 2 illustrates the average max_frequency values for each group. The error bars indicate a relatively small margin of error, suggesting that using frequency‑based grouping is a feasible approach for our fairly stable set of 117 excerpts.

Figure 2
Group‑level estimated marginal means for max_ frequency. Error bars indicate 95% confidence intervals.
While perceptually coherent, the frequency‑based taxonomy varies across singers, as corresponding frequency ranges differ individually. This makes it unsuitable as a standard ground truth for supervised learning. We therefore used it solely as a reference for interpretation and performer insight. That is, the frequency‑based taxonomy is discussed qualitatively, whereas the supervised models in the following sections use the EMVD labels as operational classes for supervised training and evaluation.
4.2 Dataset and evaluation metrics
In this work, we use the EMVD by Tailleur et al. (2024),10 consisting of 1068 audio instances (approximately 100 min of recordings) of eight different vocal techniques performed by 27 singers (including four women). Due to uneven distribution of samples across categories, our experiments focused exclusively on the four vocal techniques with at least 100 instances per class – Clear Voice, Black Shriek, Death Growl and Hardcore Scream – for a total of 946 audio instances (for its distribution, see Table 1). In the EMVD dataset, each audio recording is rated on a scale from 0 to 2 points, where 0 points indicates insufficient representation of the intended vocal technique, i.e. unsuitable for deep learning (Tailleur et al., 2024); 1 point indicates partial representation with moderate quality; and 2 points indicates excellent representation, closely matching the expected technique.
Table 1
Distribution of samples across categories (i.e. vocal techniques), considering both the EMVD and the selected samples used for the perceptual versus ML experiments. The number of samples and durations in seconds are given.
| Category | EMVD | Selected | ||
|---|---|---|---|---|
| Samples | Duration (s) | Samples | Duration (s) | |
| Clear Voice | 303 | 2367.8 | 11 | 186.7 |
| Black Shriek | 191 | 1360.7 | 20 | 321.9 |
| Death Growl | 222 | 1551.2 | 21 | 356.5 |
| Hardcore Scream | 230 | 1677.9 | 19 | 266.6 |
| Total | 946 | 6957.6 | 71 | 1155.1 |
To assess human perception and ML testing, the same audio samples were considered (from now on, referred to as the ‘test set’). Only samples awarded 2 points were selected to ensure reliability. In addition, samples from the ‘lyric’ category of the dataset were chosen for this set: Each audio sample in this category is roughly 15 s long, with lyrics chosen by the singers themselves, consistently used across all vocal techniques. The final test set included 71 samples by 23 singers. Four additional samples by different singers were used as examples during the human‑perception experiment but were not included in the test set.
To evaluate both human perception and ML classification, we use unweighted average recall (UAR), precision and recall. UAR, also known as balanced accuracy, is the average recall across all categories and is more suitable than traditional accuracy for handling imbalanced class distributions (Bekkar et al., 2013). We also present confusion matrices to illustrate confusions among categories.
4.3 Listening test
We targeted a group of participants with prior auditory experience related to extreme vocal techniques. A total of 158 participants aged from 18 to 56 years (M = 29.48, SD = 8.05), including 23 females, 123 males, five non‑binary/diverse individuals, four individuals who preferred not to disclose their gender, and three individuals who left the field blank, were included in the analysis (see Appendix 3 for the participant flow diagram). They were mainly recruited through the following channels: (1) niche music enthusiast communities, (2) metal club membership and (3) public online announcements. The procedures used adhere to the tenets of the Declaration of Helsinki.11
After signing informed consent forms guaranteeing the use of their anonymous responses for research, participants first listened to demonstration audio samples from the four vocal categories (Clear Voice, Black Shriek, Hardcore Scream and Death Growl). These four demonstration samples served as canonical examples to familiarise participants with the task and were not included in the evaluation set. They were selected from the EMVD dataset as high‑quality, prototypical instances, according to the assigned grade given in the dataset, i.e. rating = 2 points (Tailleur et al., 2024). To minimise potential speaker‑identity cues, we selected the four demo exemplars manually under three constraints: each exemplar had a rating of 2 points according to the EMVD quality grade, the four exemplars came from four different vocalists and we avoided using the same vocalist for the demo exemplar and evaluation items of the same class wherever possible.12 Participants were asked the following question for each audio clip: ‘Which vocal technique do you think is used in this clip?’ Then, they annotated 20 randomly‑assigned audio excerpts in a force‑choice task where they could select only one category. To avoid the influence of genre associations on the participants, we removed genre indicators from the voice‑type labels, as shown in Figure 3.

Figure 3
Listening‑test interface used in the human‑perception experiment. Genre labels were removed to minimise genre‑related bias.
We conducted the experiment online using the SoSciSurvey platform. A total of 71 audio excerpts were evaluated, including 11 Clear Voice, 20 Shriek, 21 Growl and 19 Scream samples. Since previous experiments demonstrated listeners’ ability to distinguish between clean voice and distorted voice quite effectively (Tailleur et al., 2024), we included fewer clean voice samples. In each round, to preserve accuracy and avoid overloading participants, only 20 randomly assigned excerpts were assessed by each individual. Stimuli were presented using a within‑subjects randomisation design implemented in SoSciSurvey such that each participant evaluated 20 randomly selected audio clips drawn from the same stimulus pool. Due to the random‑assignment procedure implemented in SoSciSurvey, individual audio clips were not presented an equal number of times; across the 71 clips, the number of annotations per item ranged from 49 to 59 (M = 53.64). Each sample was annotated by at least 48 participants. Upon completing the listening test, participants filled out demographic information, including their gender, age, preferred music genre(s), type of auditory experience with extreme vocal techniques and the number of years they had been exposed to such music. Finally, participants could share any additional thoughts or request feedback on their responses through an open‑ended field.
As shown in Figure 4, regarding auditory experience, the majority of participants (63.29%) identified themselves as ‘I am a fan/regular listener’. In addition, 28 participants (17.7%) indicated that they ‘perform extreme vocals’, while seven participants reported engaging in music production or recording related to these techniques. Notably, in terms of years of exposure (see Figure 5), 66 participants (41.8%) reported more than 10 years of experience, while only six had less than 1 year of exposure, showing the rich auditory domain experience possessed by the participants.

Figure 4
Relationship to extreme vocal techniques.

Figure 5
Familiarity with extreme vocal techniques.
In terms of music genre preferences, participants showed a wide range of interests. Figure 6 shows the percentage of participants selecting each given genre in the multi‑select genre‑preference question: ‘Which music genres that use extreme vocals do you enjoy listening to?’ (select all that apply). ‘Other’ denotes the share of participants who provided supplementary responses via a free text field. In total, 13.3% selected all five listed music genres, 6.3% not only selected all provided genres but also supplemented their answers by reporting additional preferred music styles under the ‘Other’ category and 25.9% selected two music genres.13 Death Metal was the most popular (111 participants), followed by Black Metal (102 participants) and Hardcore (90 participants). The three assessed vocal techniques are often associated with these three genres, which guarantees consistency between the listening setup and the participants’ auditory preferences.

Figure 6
Genre preferences related to extreme vocals.
4.4 Machine learning setup
In addition to the main ML classification (i.e. the baseline classification task), we conducted three auxiliary analyses: (i) a feature group importance analysis to quantify the contribution of broad ComParE feature families; (ii) a perception‑oriented human prediction task, in which an additional SVM was trained to predict listeners’ majority decisions using the same feature representation, singer‑independent three‑fold partitioning and a hyperparameter optimisation procedure as in the baseline classification task; and (iii) an ablation analysis probing the role of perturbation and noise‑related cues by comparing a model restricted to Jitter, Shimmer and harmonics‑to‑noise ratio (HNR) features with a complementary model trained on all remaining ComParE features.
4.4.1 Baseline classification task
In audio signal processing, mel frequency cepstral coefficients (MFCCs) and spectrograms have long been standard due to their alignment with human auditory perception (Purwins et al., 2019). Prior work has confirmed their effectiveness for vocal technique recognition (Stadler et al., 2023), including in metal vocal datasets such as MVD (Kalbag and Lerch, 2022) and EMVD (Tailleur et al., 2024), where log‑mel spectrograms were primarily used for convolutional neural network (CNN) input. While binary classification between Clear Voice and extreme vocals yielded high accuracy (93%), multi‑class performance declined significantly (Tailleur et al., 2024), likely due to more subtle spectral distinctions among distorted vocal styles.14 Additionally, studies have shown that extreme techniques like Death Growl and Scream exhibit higher Jitter and Shimmer, and lower HNR (Kato and Ito, 2013), suggesting that incorporating a broader set of acoustic features could further improve classification performance.
The ComParE feature set (Eyben et al., 2010) is a suitable candidate for our purposes and has shown good performance in similar studies (Parada‑Cabaleiro et al., 2023; Xu et al., 2022). It includes 6373 statistical functionals computed from 65 low‑level descriptors (LLDs) and their delta coefficients (including MFCCs, spectral, prosodic and voice‑quality features). We extracted features using the openSMILE toolkit (Eyben et al., 2010) with default parameters. To reduce feature dimensionality, computational cost and improve model performance (Guyon and Elisseeff, 2003; Kohavi and John, 1997), we use the SelectKBest feature‑selection method from scikit‑learn by Pedregosa et al. (2011), which evaluates the importance of each feature based on the ANOVA F‑test and selects the top k features.
As in similar tasks (Kalbag and Lerch, 2022; Parada‑Cabaleiro et al., 2023; Stadler et al., 2023; Xu et al., 2022), we used a SVM implemented on scikit‑learn by Pedregosa et al. (2011) with a linear kernel. We performed hyperparameter optimisation using a nested grid search within each training fold only, with singer‑wise splits applied inside the training pool, while test singers were fully held out and never used for model selection (Pedregosa et al., 2011); the candidate grids were defined a priori based on preliminary manual tuning (feature numbers [100, 200, 300, 500, 700, 900, 1000]; regularisation parameters [1, 0.1, 0.001, 0.0001, 0.00001, 0.000001]). Additionally, we conducted an exploratory feasibility check with a simple two‑layer CNN using log‑mel spectrogram input and dynamic padding in the nine‑class task (Clear Voice, Black Shriek, Death Growl, Deep Gutturals, Effect, Grind Inhale, Hardcore Scream, Pig Squeal and Tunnel Throat) but observed notably lower performance (micro accuracy 70%; macro average 32%); for parity, we also trained a nine‑class linear SVM with ComParE features and SelectKBest (top 100), achieving a micro accuracy of 77.1% and macro average of 52%. The corresponding nine‑class confusion matrices (CNN and SVM) are reported in Appendix 1, while the main analyses reported in the Results section focus on the reduced four‑class taxonomy.
To avoid the model learning individual singer characteristics and prevent data leakage, we split the data based on Singer ID and adopted a singer‑independent three‑fold evaluation scheme. Test folds were constructed from the human‑evaluation stimulus pool by assigning singers to three non‑overlapping groups using a priority‑queue (heapq) load‑balancing procedure based on the number of human‑evaluated stimuli per singer so that the total number of test stimuli was approximately balanced across folds. We carried out three independent experiments, each time on a different test set containing singers not seen in the training/validation sets (Fold 0 test singers: [2, 9, 10, 11, 12, 19, 23], 24 audio excerpts; Fold 1: [1, 3, 5, 13, 14, 16, 17], 24 audio excerpts; Fold 2: [4, 7, 8, 15, 18, 21, 22, 25], 25 audio excerpts).
Figure 7 depicts the distribution of samples across sets. To ensure comparability with the human‑perception experiment, the accumulative test sets across experiments correspond to the same 71 excerpts assessed in the perception experiment. Groups with more excerpts were merged with those with fewer excerpts to ensure that test excerpts per experiment were within ±2. Table 2 shows the distribution of excerpts across classes in the test set for each experiment. When discussing the ML results (see Section 5.2), we report the mean across the UAR obtained from the three experiments.

Figure 7
Sample distribution across the three independent experiments. Each experiment presents a different test set (together, totalling the 71 samples used in the perceptual study). For training and evaluation sets, is given across the singer‑independent three‑fold group cross‑validation (by singer ID) performed for optimisation.
Table 2
Distribution of the 71 samples included in the test set across vocal techniques. To guarantee a singer‑independent ML setup, the samples were distributed across three independent experiments with unique singers in the test set.
| Technique | Experiment 1 | Experiment 2 | Experiment 3 |
|---|---|---|---|
| Clear Voice | 4 | 4 | 3 |
| Black Shriek | 6 | 6 | 8 |
| Death Growl | 5 | 9 | 7 |
| Hardcore Scream | 9 | 3 | 7 |
| Total | 24 | 22 | 25 |
4.4.2 Feature group importance analysis
To assess which broad acoustic cue families contribute most strongly to the classifier, we conducted a feature group importance analysis based on the linear SVM model. After standardising all ComParE features, we used the mean absolute SVM weights as an indicator of feature contribution. Individual features were grouped into six broad families (MFCC, Spectral, Shimmer, Jitter, HNR and a residual ‘Other’ group).15 For each fold, absolute weights were aggregated within each feature group and mean group importance as well as within‑model ranks, which were computed across folds. This procedure allows us to compare which feature families receive the strongest weights in the classifier and to analyse differences in cue weighting across classes.
4.4.3 Human‑prediction task
In order to link the classification task more directly to the perception experiment, we trained an additional linear SVM that predicts listener majority decisions. For each of the 71 audio excerpts in the listening test, we derived a majority category over all valid forced choice responses (ClearVoice, Shriek, Growl and Scream). The singer‑independent three‑fold split was predefined based on singer ID and identical to the split used for the main ML classifier. Accordingly, the test partition in each fold was exactly the same as in the baseline classification task. For the human‑prediction task, we kept these fold‑specific test items fixed and replaced the corpus labels with the listeners’ majority decision labels for evaluation. Using the same ComParE feature representation and the same hyperparameter‑optimisation procedure as in the baseline classification task, we fitted a human‑prediction SVM on the fold‑specific test items for which a valid majority label was available and evaluated it with UAR and the precision recall area under the receiver‑operating characteristic curve (PR‑AUC).
4.4.4 Feature group importance comparison
Feature group importance scores and per‑class group scores were also computed for the human‑prediction task (see Section 4.4.3), following the same procedure as carried out for the baseline model (see Section 4.4.2). Therefore, it was possible to directly compare which ComParE feature families receive the strongest weights when approximating listener majority decisions.
4.4.5 Ablation analysis
As outlined above, Jitter, Shimmer and HNR have been reported to play an important role in Growl and Scream (Kato and Ito, 2013). Based on this, we formulated a hypothesis regarding the source of confusion between Shriek and Scream. We hypothesised that, although these two categories share similar overall acoustic cue profiles, the baseline classification model may distinguish them in part through finer‑grained combinations within perturbation‑related cue families rather than through differences between broad feature groups.
To evaluate this hypothesis, we conducted additional analyses focusing on two models: one restricted to Jitter, Shimmer and HNR features and another complementary one trained on the remaining ComParE feature set after removing these features. We compared the results from the baseline task (full feature model) with those from the restricted model (encompassing only Jitter, Shimmer and HNR features) and with the complementary model (encompassing all ComParE features except Jitter, Shimmer and HNR). We then examined the bidirectional confusion between Shriek and Scream under each feature configuration.
5. Results
In the following, we report the results obtained from the human‑perception evaluation and from the ML experiments. For readability, from this point onward, we consistently use the short label forms in the running text (e.g. Shriek instead of Black Shriek).
5.1 Human‑perception experiment
Figure 8 shows the confusion matrix for the human‑perception experiment (UAR of 76.2%). Clear Voice achieved the highest recall (97%) and Growl reached nearly 80%. In contrast, the recall for Shriek and Scream were lower, both below 70% (68.6% and 60.9%, respectively). The precision scores for these two techniques were also relatively low, especially for Shriek (63%), even below Scream (68%), as the former attracted the highest confusion from the other classes (see darker column for Shriek in the confusion matrix reported in Figure 8).16

Figure 8
Human perception. Confusion matrix for the identification of the four extreme vocalisations.
All in all, these results indicate that Clear Voice and Growl were more reliably identified by human listeners, while Scream, but in particular Shriek, were more often confused with other categories.
Although Growl and Shriek show some differences in their spectrogram representations (Growl is more concentrated in low frequencies, whereas Shriek is more concentrated in high frequencies), 15.2% of Growl samples were still misclassified as Shriek. The classification performance for Scream was the worst, with the highest confusion with Shriek (23.5%), which may be attributed to the similar frequency characteristics of Scream and Shriek (see spectrograms in Figure 1). One possible explanation for this is that listener expectations and mental‑categorisation strategies influenced their decisions: When exposed to a series of extreme vocal styles, participants may have second‑guessed certain Clean Voice samples, interpreting subtle vocal effort or coloration as indicative of a more distorted technique. This suggests that perception was shaped not only by acoustic features but also by the context in which the samples were evaluated. Moreover, it is possible that some listeners interpreted the Shriek category as a sort of fallback label, assigning uncertain samples there when no clear match was perceived. Such tendencies may indicate that Shriek was perceived as less acoustically distinct compared to other categories.
5.2 Machine learning experiments
5.2.1 Baseline classification task
We conducted three independent ML experiments using a singer‑independent three‑fold group cross‑validation scheme (by singer ID) for optimisation. In each fold, the test singers were fully held out and differed across folds. The reported results are averaged across the three folds. The SVM demonstrates good performance, achieving a UAR ( in %) of . Overall, the pooled test‑set UAR was 90.0% (95% bootstrap CI: [82.9%, 96.3%]; N = 71). The confusion matrix (see Figure 9) illustrates the classification results and confusion patterns. The classification performance across categories was high, with Clear Voice reaching both a precision and recall of 100%. Growl achieved a recall of 85.7% and precision of 89.1%. For Shriek and Scream, even if varying in recall (85.0% and 89.5%, respectively), precision scores were both 85.6%.

Figure 9
ML. Confusion matrix for the classification of four extreme vocalisations.
In addition, per‑class PR‑AUC values were: Clear Voice , Shriek , Growl and Scream (macro avg. ; micro avg. ). The combination of high precision and recall across all categories suggests that the model is capable of identifying each vocal technique both accurately and reliably, thus outperforming humans not only in overall accuracy but also in consistency across categories.
Clear Voice was classified perfectly, which also confirms previous findings (Tailleur et al., 2024) that distorted and non‑distorted vocals can be effectively recognised by ML models. Shriek achieved a recall of 85%, with 15% of samples misclassified as Scream. This result is consistent with the human‑perception experiment, where 22.4% Shriek samples were misclassified as Scream. The ML model reproduces two salient human‑confusion directions (see Figure 8 versus Figure 9 for comparison). Growl is frequently misclassified as Shriek (ML model 14.3%; human 15.2%). Scream is also misclassified as Growl at a comparable rate (ML model 10.5%; human 11.5%). These alignments suggest that the main errors reflect shared perceptual ambiguities between these technique pairs.
Most notably, human participants frequently confused Scream and Shriek, with Scream being misclassified as Shriek in 23.5% of cases. In contrast, such confusion did not exist in the ML classification results. This may be attributed to the model’s ability to perform fine‑grained feature extraction, using a greater number of acoustic features to detect subtle differences that are often imperceptible to human listeners. Moreover, ML models are not subject to subjective biases and can consistently evaluate clarity and frequency distribution.
5.2.2 Feature group importance analysis
Across folds, the highest mean group weight observed was for Shimmer, followed by Other, HNR, Spectral and Jitter, with MFCCs contributing the least (see Figure 10). At the class level (mean standard deviation), Shriek and Scream showed highly similar profiles, with the highest mean group weights in Shimmer and Jitter, whereas Clear Voice was driven by Spectral and MFCC features. Growl was most strongly associated with Shimmer and showed elevated contributions from Other and HNR, indicating that Growl, unlike the other categories, is mainly characterised by stronger amplitude‑perturbation cues, while Shriek and Scream likely depend on finer‑grained patterns within these perturbation groups.

Figure 10
Per‑class broad ComParE feature group importance for the EMVD label model (mean standard deviation across folds), with distribution shown within class ranks.
In line with vocology accounts that extreme vocal techniques arise from non‑traditional vibratory behaviours in the vocal folds and other laryngeal structures (Erbe, 2014; Nieto, 2008; Smialek et al., 2012), the prominence of perturbation‑related groups in our model suggests that irregular vibration and increased noise components provide salient cues for distorted techniques. This is consistent with Kato and Ito (2013), who reported higher Jitter and Shimmer and lower HNR in Growl and screaming voices compared to Clear Voice, highlighting perturbation‑ and harmonicity‑related differences between distorted and non‑distorted phonation.
5.2.3 Human‑prediction task
We trained an additional ComParE SVM to predict human majority decisions. Using the same experimental setup as in the baseline classification task, the model achieved a pooled test performance of 76.1% accuracy and 78.3% UAR (N = 71). Across folds, the highest mean group weight in the human‑prediction model was observed for the MFCC category, followed by HNR, Shimmer, Jitter and Other, with Spectral contributing least (see Figure 11). At the class level (mean ± standard deviation), Shriek was most strongly associated with MFCC and HNR, Clear Voice showed highest contributions from HNR, Growl showed highest contributions from MFCC and Shimmer and Scream relied comparatively more on Jitter and HNR.

Figure 11
Per‑class broad ComParE feature group importance for the human‑prediction model based on perception labels (mean standard deviation across folds), with distribution shown within class ranks.
5.2.4 Feature group importance comparison
Figure 12 compares the broad feature group profiles between the baseline classification model and the human‑prediction model. While the baseline classification model is dominated by Shimmer (rank 1) and places the MFCC category in rank 6, the human‑prediction model shifts the MFCC category to the top (rank 1) and pushes Spectral to rank 6. This ranking shift in our human‑prediction task suggests that listener majority decisions may place greater weight on MFCC‑based timbral cues, supporting the interpretation that the human prediction model relies more on perceptually salient timbral information and that humans may also place more weight on HNR‑related cues linked to perceived periodicity versus noisiness and voice quality. In contrast, the dataset label classifier is dominated by shimmer‑related cues; relative to the main model, Spectral ranks lowest in the human‑prediction model, and Other features become less prominent in this comparison.

Figure 12
Comparison of broad ComParE feature group importance between the baseline classification model and the human‑prediction model, reported as mean ± standard deviation across folds with within‑model ranks.
5.2.5 Ablation analysis
We evaluate the hypothesis that the baseline classification model may distinguish Shriek and Scream through finer‑grained combinations within perturbation‑related cue families to a greater extent than through differences between broad feature groups. In our feature group analysis (see Figure 10), Shriek and Scream indeed showed highly similar group level importance profiles, both being dominated by perturbation‑related cues, especially Shimmer and Jitter. By contrast, Growl showed a more distinct pattern, with stronger Shimmer and additional contributions from HNR and Other. Moreover, the feature group analysis in both the human‑prediction task and the baseline classification task showed that HNR ranked second and third, respectively (see Figure 12).
The ablation results further support this interpretation. When only Jitter, Shimmer and HNR features (i.e. the restricted model) were retained, pooled UAR decreased from 90.0% to 76.7% and bidirectional Shriek or Scream confusion increased: Shriek to Scream remained 15.0%, while Scream to Shriek increased from 0.0% to 10.5%; see Figure 13. When Jitter, Shimmer and HNR features were removed (i.e. the complementary model), pooled UAR also decreased to 82.3%, and the opposite confusion direction increased further: Shriek to Scream decreased from 15.0% to 10.0%, while Scream to Shriek increased to 15.8%; see Figure 14. Taken together, these results support the view that the Shriek versus Scream distinction is not primarily based on different broad acoustic cue families but more likely on finer‑grained combinations within and across cue groups, which may also help explain the elevated human confusion between the two categories.

Figure 13
Ablation analysis: restricted model (based on Jitter, Shimmer and HNR). Confusion matrix.

Figure 14
Ablation analysis: complementary model (ComParE excluding Jitter, Shimmer and HNR).
6. Discussion
From the results presented, ML classification clearly outperforms human perceptual performance, particularly for Shriek and Scream, where listeners show elevated confusion. We interpret this human versus ML discrepancy by jointly considering the ablation study and the human‑prediction model. Together, these analyses suggest that voice quality‑related cues, including Jitter, Shimmer and HNR, are salient in both label spaces but not sufficient on their own, and that robust category separation depends on finer‑grained combinations within and across feature groups. At the same time, the contrast between the EMVD‑based classifier and the human‑prediction model indicates that humans and the classifier do not weight these cues in the same way. Part of the human versus ML gap may therefore reflect different cue prioritisation under decontextualised listening, rather than limited access to any one feature group alone.
For human listeners, distinguishing between Shriek and Scream is particularly challenging due to their similar frequency ranges and overlapping high‑frequency noise components. Without professional vocal training, it is difficult to accurately perceive and differentiate these timbral nuances. In addition, human perception tends to rely more heavily on familiarity with the music, contextual cues (Martin, 1999) and emotional expression (Juslin and Västfjäll, 2008), rather than purely acoustic characteristics.
In our experiment, we employed isolated dry vocals without musical accompaniment, which might have penalised the listeners while facilitating the ML model.17 This asymmetry is expected because human judgements are typically anchored in musical framing and expressive cues, whereas the classifier operates on decontextualised acoustic regularities. Consequently, the human–ML gap observed should be interpreted as partially contingent on the dry‑vocal condition rather than as a direct proxy for in‑the‑wild listening.
Listener judgements can be influenced by familiarity with commercial recordings and stylistic conventions. At the same time, copyright constraints limit the systematic use and public release of fully contextualised commercial tracks, which is one reason we relied on isolated vocal excerpts. Our interviews further suggest that vocal technique labels are strongly correlated with musical style and production context in practice, such that listeners may partly rely on stylistic framing when making technique judgements. This linkage between vocal performance techniques and stylistic conventions is also discussed in Kennedy’s dissertation (see Kennedy (2018) Chapter 3, ‘Vocals’, pp. 69–79). Indeed, many participants reported that the absence of musical context made the classification task challenging to a certain extent. This seems to contradict prior research suggesting that expert listeners, through long‑term exposure to such vocal techniques, develop the ability to focus their attention on the most informative acoustic cues, a perceptual skill highly generalisable across unfamiliar stimuli (Olsen et al., 2018).
To evaluate potential bias from our controlled setup, future studies should include musical context as a factorial variable (unaccompanied versus accompanied × vocal class) in human versus ML comparisons. External validation of independent material will also be important for assessing the generalisability of the present findings more fully. Several participants noted in the open‑ended comments that, although vocal techniques play a critical role in this type of music, they are often overlooked or insufficiently recognised by general listeners. According to these participants, vocal techniques tend to receive less attention compared to instrumental elements, despite their central expressive function in extreme metal. Indeed, the fact that many listeners, despite long‑term exposure to this style, were unable to accurately identify the specific vocal techniques they encountered, not only reflects a need for further educational outreach and academic exploration of these specialised vocal styles but also highlights the lack of a standardised taxonomy, which may contribute to confusion and inconsistent terminology. Addressing this gap is essential for both listener understanding and the development of reliable classification systems, as explored in the first phase of this study.
Such taxonomies could also benefit educational practices, by providing clearer frameworks for teaching and learning the nuances of extreme vocal performance. This observation aligns with broader structural critiques in music genre recognition (MGR) research. For instance, Green et al. (2024) argue that many classification systems rely heavily on oversimplified genre labels and rarely consider vocal technique as a central stylistic marker, as their primary goal is improving accuracy and other performance metrics while disregarding the cultural and aesthetic functions of vocal performance.
Furthermore, genre is not a static set of sonic features but a dynamic, socially constructed category, shaped by the practices of musicians, listeners, critics and institutions. Indeed, institutional understandings of genre vary widely, and vocal techniques (particularly in extreme or non‑Western styles) are often neglected. As shown by Green et al. (2024), among 560 works on MGR published between 2013 and 2022, only 61 addressed music from non‑Western traditions. From this perspective, improving tagging and retrieval systems based on specific vocal techniques could enhance recommendation relevance and accuracy while also providing a richer pedagogical and research‑oriented framework. Given the global diffusion and cultural diversity of extreme vocal practices, classification frameworks should better accommodate such multiplicity and avoid homogenising distinct styles.
The classification systems commonly employed in both streaming platforms and academic research predominantly rely on broad genre categories such as ‘metal’ and ‘rock’. The genre taxonomy presented to users may also be considerably simpler than the more fine‑grained internal classification structures used for clustering and recommendation. This is particularly evident on platforms like Spotify, where the user‑facing genre taxonomy appears markedly restricted when compared to the extensive internal genre hierarchies maintained by these systems (Krogh, 2025). Even genres widely recognised for their stylistic diversity and subcultural specificity, such as metal, are rarely subdivided in user‑facing interfaces. Instead of offering curated selections based on distinct subgenres like thrash, death, doom or black metal, Spotify typically features generic, chart‑driven playlists such as ‘Popular Metal’ or ‘Metal Heavyweights’. This lack of detailed subcategorisation also obscures the diversity of associated vocal techniques, which are often deeply embedded within specific subgenres. However, such services do not fully disclose the underlying taxonomies on which their recommendation algorithms operate, which may extend beyond the representative labels presented to the public.
Given that extreme vocal techniques are closely tied to particular stylistic and acoustic conventions, such oversimplified classification may hinder both user navigation and scholarly analysis. Since extreme vocal techniques are inherently embedded within specific musical subgenres (each characterised by distinct stylistic and acoustic conventions), current classification systems often fail to reflect these nuances. As a result, while direct searches may yield relevant results, subsequent recommendation or autoplay functions often lead to stylistically divergent genres. For example, after listening to a track labelled as ‘old school metal’, the system may automatically suggest or play songs from a fundamentally different genre such as metalcore. This mismatch highlights a deeper structural issue: existing taxonomies are insufficiently aligned with the sonic and stylistic features that define subgenre distinctions, particularly in metal and related genres.
To address these limitations, future tagging and retrieval systems could incorporate vocal technique distinctions as an organising principle. Doing so could improve both the accuracy and relevance of content recommendations, particularly in genres where vocal style functions as a key aesthetic marker. Moreover, given that extreme vocal practices, such as globally diverse singing styles and regionally specific growling techniques, are used across a wide range of global traditions and that the niche musical cultures employing such techniques also have dedicated followings worldwide, classification frameworks should evolve to better reflect this stylistic and cultural diversity.
Indeed, one participant in the test suggested in the final open comments to incorporate additional vocal traditions, such as those found in jazz or folk‑based singing, which could further improve the generalisability of the model and support more inclusive classification strategies. That said, the current study has limitations. On the one hand, the controlled conditions of our experimental setup (based on dry vocals), may have introduced specific biases favouring ML performance while disadvantaging human perception. On the other, some participants noted that the dataset could benefit from greater variation in vocal registers, technique types and performer demographics. Thus, future research should investigate more realistic scenarios as well as expand existing datasets in the mentioned directions, promoting model generalisability and enabling more inclusive and representative approaches to vocal classification.
7. Conclusion
In summary, this study presented a comparative evaluation of human and machine classification performance on extreme vocal techniques, using the EMVD dataset. By analysing perceptual judgements from 158 expert participants alongside ML outputs from an SVM model trained on ComParE features, we demonstrated that automatic classification can outperform human perception under the present experimental conditions, particularly for isolated unaccompanied vocal excerpts, achieving an average UAR of 90% compared to 76.2% for human listeners. These findings suggest that data‑driven methods hold strong potential for supporting the annotation and retrieval of under‑represented vocal styles in MIR systems. At the same time, perceptual confusion patterns point to areas where human intuition still carries critical value. Our additional feature group, ablation and human‑prediction analyses further suggest that the human–machine gap reflects not only differences in the use of acoustic information but also differences in cue prioritisation: human perception relies more on MFCC and HNR and is shaped by familiarity, context and training, whereas machine classification appears to rely on finer‑grained combinations of features within and across perturbation related cue groups. Future work could aim to further integrate praxis‑based and perceptually‑informed approaches into computational modelling, explicitly examine context effects by comparing unaccompanied and accompanied vocal conditions, explore feature sets tailored specifically to extreme vocals and expand to diverse vocal traditions beyond Western metal to foster a more inclusive understanding of vocal expressivity.
Acknowledgements
We sincerely thank all the participants and the vocalists who took part in our work. We would also like to thank the anonymous reviewers who contributed to bringing this work to its current form. We are further grateful to Paul Peters and Luca Matsukawa for their assistance in preparing the stimulus material. In the methodological notes, they are referred to as Member A and Member B, respectively. We also thank Marcus Erbe, whose course provided important inspiration for this study. We are also grateful to Philippa Ovenden for her help with proofreading the manuscript.
Data Accessibility
The full interview transcripts are provided in Appendix 2.
The audio excerpts used for the taxonomy validation in Section 4.1 were drawn from 117 commercial recordings provided in the context of the course from which the initial study design emerged. Due to copyright restrictions, these commercial audio excerpts cannot be redistributed as part of the present publication. All subsequent empirical and computational analyses, including the human listening experiment and the ML experiments, were conducted using the publicly available Extreme Metal Vocals Dataset: https://zenodo.org/records/8406322.
The code repository is available at: https://github.com/extremevocalclassifier/extremevocalclassifier.
Ethics and Consent
This study involved a listening experiment and interviews with adult participants. All participants provided informed consent before taking part in the study. Participation was voluntary, and participants could withdraw from the study at any time. Listening test responses were analysed anonymously and only in aggregated form. Interviewees provided consent for their statements to be used in the present research. Where applicable, identifying details are reported only with the interviewees’ permission.
At the time the study was conducted, the ethics committee of the Faculty of Arts and Humanities, University of Cologne, had not yet been established. The committee was therefore consulted retrospectively after the completion of the study. On the basis of this retrospective consultation, the committee indicated that the study did not raise further ethical concerns, provided that all participants were adults and that the data were processed anonymously.
Competing Interests
The authors have no competing interests to declare.
Authors’ Contributions
Xuhong Qiu conducted the main empirical, interview‑based and computational work of the study and prepared the manuscript. Emilia Parada‑Cabaleiro proposed the initial human–machine comparison framework, supervised the project and contributed to revising and editing the manuscript. Both authors approved the final version of the manuscript.
Notes
[1] In this study, EMVD labels are used as operational dataset categories for analysis, rather than as universally fixed terms, since terminology may vary across scenes and practitioner communities. We retain the EMVD label ‘Clear Voice’ for consistency with the dataset, listening test and ML pipeline. In everyday usage (including our interviews), participants more often used the term ‘Clean Voice’, typically in contrast to ‘Harsh Voice’ or ‘Distorted Voice’.
[2] The code repository is available at: https://github.com/extremevocalclassifier/extremevocalclassifier.
[3] In their paper, Tailleur et al. used the term ‘heavy metal’ to describe the scope of their study. However, ‘heavy metal’ can refer either to a historically situated style associated with 1980s metal or, from an outsider perspective, to metal music more broadly. To avoid this ambiguity, we refer to the relevant repertoire as ‘extreme metal’ in this paper for greater terminological precision.
[4] For example, in the Swedish band Merciless, we can hear two clearly different vocal styles between their first demo ‘Behind the Black Door’ and the album ‘The Awakening’. Although the demo is evidently not yet fully mature, the difference in vocal style is still clearly audible.
[5] The first author conducted the interviews and Python analyses. Member A prepared the excerpts by segmenting commercial recordings from course materials. The statistical testing reported here was initially performed by Member B and independently verified by the second author. All subsequent experiments were conducted solely by the two authors; other research group members contributed only to the excerpt preparation and the initial validation reported in this section.
[6] The 117 audio excerpts are drawn from commercial music recordings and were provided to the research group as part of course materials by the course instructor.
[7] Member A manually segmented the commercial recordings into excerpts; Member B used SPSS automatic binning of max_frequency to assign excerpts to the low‑, mid‑ and high‑frequency groups.
[8] Prior to the FFT, we applied a Hann window to the full‑length waveform to reduce spectral leakage due to boundary discontinuities.
[9] The spectral centroid validation and the associated SPSS analyses reported in this paragraph were performed by the first author.
[10] The dataset is available at: https://zenodo.org/records/8406322.
[12] The specific files used were: Singer7_ClearVoice_Mid_lyrics for Clear Voice, Singer5_BlackShriek_High_lyrics for Black Shriek, Singer9_ DeathGrowl_Low_lyrics for Death Growl and Singer8_Hardcore Scream_High_lyrics for Hardcore Scream. Although the file names indicate different frequency ranges, this is determined by the characteristics of the respective vocal techniques (Tailleur et al., 2024). See Section 4.2 for a detailed description of the techniques.
[13] We included an ‘Other’ option because, in the pilot test, participants repeatedly noted that the listed genres were not exhaustive and because the ‘Other’ field allows such additions to be collected in a format that is easy to aggregate for analysis.
[14] Tailleur et al., reported 93% micro and macro accuracy for binary classification between Clear Voice and distorted vocals, and 75% micro accuracy and 70% macro accuracy for four class classification using EfficientNet with log‑mel spectrogram input. Because the exact experimental setup is not fully identical, we do not treat this as a strict one‑to‑one benchmark comparison, but it nevertheless provides a useful point of reference for the present four class task.
Additional File
The additional file for this article can be found as follows:
