
Figure 1
Log‑mel spectrograms of the four investigated vocal techniques (Clear Voice, Black Shriek, Death Growl and Hardcore Scream) performed by Singer ID 10 (male) for ‘Mid_a’ (vowel [a] in a mid‑frequency range). Each panel shows the first 4 s of the recording. The x‑axis shows time (s) and the y‑axis shows the mel frequency with tick labels in Hz. The white box indicates the main band (20th–90th percentile energy interval).

Figure 2
Group‑level estimated marginal means for max_ frequency. Error bars indicate 95% confidence intervals.
Table 1
Distribution of samples across categories (i.e. vocal techniques), considering both the EMVD and the selected samples used for the perceptual versus ML experiments. The number of samples and durations in seconds are given.
| Category | EMVD | Selected | ||
|---|---|---|---|---|
| Samples | Duration (s) | Samples | Duration (s) | |
| Clear Voice | 303 | 2367.8 | 11 | 186.7 |
| Black Shriek | 191 | 1360.7 | 20 | 321.9 |
| Death Growl | 222 | 1551.2 | 21 | 356.5 |
| Hardcore Scream | 230 | 1677.9 | 19 | 266.6 |
| Total | 946 | 6957.6 | 71 | 1155.1 |

Figure 3
Listening‑test interface used in the human‑perception experiment. Genre labels were removed to minimise genre‑related bias.

Figure 4
Relationship to extreme vocal techniques.

Figure 5
Familiarity with extreme vocal techniques.

Figure 6
Genre preferences related to extreme vocals.

Figure 7
Sample distribution across the three independent experiments. Each experiment presents a different test set (together, totalling the 71 samples used in the perceptual study). For training and evaluation sets, is given across the singer‑independent three‑fold group cross‑validation (by singer ID) performed for optimisation.
Table 2
Distribution of the 71 samples included in the test set across vocal techniques. To guarantee a singer‑independent ML setup, the samples were distributed across three independent experiments with unique singers in the test set.
| Technique | Experiment 1 | Experiment 2 | Experiment 3 |
|---|---|---|---|
| Clear Voice | 4 | 4 | 3 |
| Black Shriek | 6 | 6 | 8 |
| Death Growl | 5 | 9 | 7 |
| Hardcore Scream | 9 | 3 | 7 |
| Total | 24 | 22 | 25 |

Figure 8
Human perception. Confusion matrix for the identification of the four extreme vocalisations.

Figure 9
ML. Confusion matrix for the classification of four extreme vocalisations.

Figure 10
Per‑class broad ComParE feature group importance for the EMVD label model (mean standard deviation across folds), with distribution shown within class ranks.

Figure 11
Per‑class broad ComParE feature group importance for the human‑prediction model based on perception labels (mean standard deviation across folds), with distribution shown within class ranks.

Figure 12
Comparison of broad ComParE feature group importance between the baseline classification model and the human‑prediction model, reported as mean ± standard deviation across folds with within‑model ranks.

Figure 13
Ablation analysis: restricted model (based on Jitter, Shimmer and HNR). Confusion matrix.

Figure 14
Ablation analysis: complementary model (ComParE excluding Jitter, Shimmer and HNR).
