Skip to main content
Have a personal or library account? Click to login
Distorted Realities: Classifying Extreme Vocals Between Harmony and Noise – A Machine–Human Evaluation of Vocal Confusion Patterns Cover

Distorted Realities: Classifying Extreme Vocals Between Harmony and Noise – A Machine–Human Evaluation of Vocal Confusion Patterns

Open Access
|Jul 2026

References

  1. Aaen, M., Sadolin, C., and McGlashan, J. (2025). Toward a taxonomy for supraglottic structure vibrations in voice: A narrative literature review. Perspectives of the ASHA Special Interest Groups, 121. 10.1044/2024_PERSP-24-00140.
  2. Bekkar, M., Djemaa, H. K., and Alitouche, T. A. (2013). Evaluation measures for models assessment over imbalanced data sets. Journal of Information Engineering and Applications, 3, 2738. 10.5121/ijdkp.2013.3402.
  3. Brackett, D. (2016). Categorizing Sound: Genre and Twentieth‑Century Popular Music. University of California Press. 10.1525/california/9780520248717.001.0001.
  4. Caffier, P. P., Ibrahim Nasr, A., and Ropero Rendon, M. D. M. (2018). Common vocal effects and partial glottal vibration in professional nonclassical singers. Journal of Voice, 32(3), 340346. 10.1016/j.jvoice.2017.06.009.
  5. Eckers, C., Hütz, D., Kob, M., Murphy, P. J., Houben, D., and Lehnert, B. (2009). Voice production in death metal singers. In Proceedings of NAG/DAGA 2009, Rotterdam, The Netherlands, pp. 17471750.
  6. Erbe, M. (2014). By demons be driven? Scanning “monstrous” voices. In In Hardcore, Punk, and Other Junk: Aggressive Sounds in Contemporary Music (pp. 5171). 10.5771/9780739176061-51.
  7. Eyben, F., Wöllmer, M., and Schuller, B. (2010). OpenSMILE – The Munich versatile and fast open‑source audio feature extractor. In Proceedings of the ACM International Conference on Multimedia (MM 2010), Firenze, Italy, pp. 14591462. 10.1145/1873951.1874246.
  8. Fuks, L., Hammarberg, B., and Sundberg, J. (1998). A self‑sustained vocal ventricular phonation mode: Acoustical, aerodynamic and glottographic evidences. KTH TMH‑QPSR, 3, 4959.
  9. Gentilucci, M., Ardaillon, L., and Liuni, M. (2018). Vocal distortion and real‑time processing of roughness. In Proceedings of the International Computer Music Conference (ICMC 2018). Seoul, South Korea.
  10. Green, O., Sturm, B., Born, G., and Wald‑Fuhrmann, M. (2024). A critical survey of research in music genre recognition. In Proceedings of the 25th International Society for Music Information Retrieval Conference (ISMIR), San Francisco, CA, USA, pp. 745782.
  11. Guyon, I., and Elisseeff, A. (2003). An introduction to variable and feature selection. The Journal of Machine Learning Research, 3, 11571182. 10.1162/153244303322753616.
  12. Guzman, M., Acevedo, K., Leiva, F., Ortiz, V., Hormazabal, N., and Quezada, C. (2018). Aerodynamic characteristics of growl voice and reinforced falsetto in metal singing. Journal of Voice, 32(6), 674683. 10.1016/j.jvoice.2018.04.022.
  13. Ishi, C. T., Sakakibara, K.‑I., Ishiguro, H., and Hagita, N. (2007). A method for automatic detection of vocal fry. IEEE Transactions on Audio, Speech, and Language Processing, 15(1), 346355. 10.1109/TASL.2007.910791.
  14. Jothimani, S., and Premalatha, K. (2022). MFF‑SAUG: Multi feature fusion with spectrogram augmentation of speech emotion recognition using convolution neural network. Chaos, Solitons & Fractals, 162, 112512. 10.1016/j.chaos.2022.112512.
  15. Juslin, P. N., and Västfjäll, D. (2008). Emotional responses to music: The need to consider underlying mechanisms. Behavioral and Brain Sciences, 31(5), 559575. 10.1017/s0140525x08005293.
  16. Kalbag, V., and Lerch, A. (2022). Scream detection in heavy metal music. arXiv preprint arXiv:2205.05580
  17. Kato, K., and Ito, A. (2013). Acoustic features and auditory impressions of death growl and screaming voice. In Proceedings of the 2013 Ninth International Conference on Intelligent Information Hiding and Multimedia Signal Processing, Beijing, China, pp. 460463. 10.1109/IIH-MSP.2013.120.
  18. Kennedy, L. (2018). Functions of Genre in Metal and Hardcore Music [PhD thesis]. The University of Hull.
  19. Kohavi, R., and John, G. H. (1997). Wrappers for feature subset selection. Artificial Intelligence, 97(1), 273324. 10.1016/s0004-3702(97)00043-x.
  20. Krogh, M. (2025). Rampant abstraction as a strategy of singularization: Genre on Spotify. Cultural Sociology, 19(1), 89107. 10.1177/17499755231172828.
  21. Li, H., Li, J., Liu, H., Liu, T., Chen, Q., and You, X. (2024). MelTrans: Mel‑spectrogram relationship learning for speech emotion recognition via transformers. Sensors, 24(17). 10.3390/s24175506.
  22. Lindestad, P. A., Södersten, M., Merker, B., and Granqvist, S. (2001). Voice source characteristics in Mongolian “throat singing” studied with high‑speed imaging technique, acoustic spectra, and inverse filtering. Journal of Voice, 15(1), 7885. 10.1016/s0892-1997(01)00008-x.
  23. , F. M. (2012). Teaching singing and technology. In K. S. Basa (Ed.), Aspects of Singing II: Unity in Understanding – Diversity in Aesthetics (pp. 88109). BDG e.V.
  24. Martin, K. D. (1999). Sound‑Source Recognition: A Theory and Computational Model. [PhD thesis]. Massachusetts Institute of Technology.
  25. McCoy, S. (2014). Singing pedagogy in the twenty‑first century: A look toward the future. In S. D. Harrison and J. O’Bryan (Eds.), Teaching Singing in the 21st Century (pp. 1320). Springer. 10.1007/978-94-017-8851-9_2.
  26. McFee, B., Raffel, C., Liang, D., Ellis, D. P., McVicar, M., Battenberg, E., and Nieto, O. (2015). librosa: Audio and music signal analysis in Python. SciPy, 1824. 10.25080/Majora-7b98e3ed-003.
  27. Nandwana, M. K., Ziaei, A., and Hansen, J. H. L. (2015). Robust unsupervised detection of human screams in noisy acoustic environments. In Proceedings of the 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), South Brisbane, QLD, Australia, pp. 161165. 10.1109/ICASSP.2015.7177952.
  28. Nieto, O. (2008). Voice transformations for extreme vocal effects. Master’s thesis. Pompeu Fabra University.
  29. Nieto, O. (2013). Unsupervised clustering of extreme vocal effects. In Proceedings of the 10th International Conference on Advances in Quantitative Laryngology, Voice and Speech Research, New York, NY, USA.
  30. Olsen, K. N., Thompson, W. F., and Giblin, I. (2018). Listener expertise enhances intelligibility of vocalizations in death metal music. Music Perception, 35(5), 527539. 10.1525/mp.2018.35.5.527.
  31. Parada‑Cabaleiro, E., Batliner, A., Schmitt, M., Schedl, M., Costantini, G., and Schuller, B. W. (2023). Perception and classification of emotions in nonsense speech: Humans versus machines. PLOS ONE, 18, 126. 10.1371/journal.pone.0281079.
  32. Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E. (2011). Scikit‑learn: Machine learning in Python. Journal of Machine Learning Research, 12, 28252830. http://jmlr.org/papers/v12/pedregosa11a.html, 10.48550/arXiv.1201.0490.
  33. Pohjalainen, J., Alku, P., and Kinnunen, T. (2011). Shout detection in noise. In Proceedings of the 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) Prague, Czech Republic, pp. 49604963. 10.1109/ICASSP.2011.5947471.
  34. Purwins, H., Li, B., Virtanen, T., Schlüter, J., Chang, S., and Sainath, T. N. (2019). Deep learning for audio signal processing. IEEE Journal of Selected Topics in Signal Processing, 13(2), 206219. 10.1109/jstsp.2019.2908700.
  35. Sadolin, C. (2012). Complete Vocal Technique. CVI Publications.
  36. Sakakibara, K.‑I., Fuks, L., Imagawa, H., and Tayama, N. (2004). Growl voice in ethnic and pop styles. In Proceedings of the International Symposium on Musical Acoustics (ISMA 2004), Nara, Japan.
  37. Smialek, E., Depalle, P., and Brackett, D. (2012). A spectrographic analysis of vocal techniques in extreme metal for musicological analysis. In Proceedings of the ICMC 2012 – International Computer Music Conference, Montréal, QC, Canada, pp. 8892.
  38. Stadler, A., Parada‑Cabaleiro, E., and Schedl, M. (2023). Towards potential applications of machine learning in computer‑assisted vocal training. In Proceedings of the International Symposium on Computer Music Multidisciplinary Research (CMMR), Tokyo, Japan, pp. 430441.
  39. Tailleur, M., Pinquier, J., Millot, L., Vogel, C., and Lagrange, M. (2024). EMVD dataset: A dataset of extreme vocal distortion techniques used in heavy metal. In Proceedings of the International Conference on Content‑Based Multimedia Indexing (CBMI), Reykjavik, Iceland, pp. 15. 10.1109/CBMI62980.2024.10859205.
  40. Titze, I. (2008). Nonlinear source‑filter coupling in phonation: Theory. The Journal of the Acoustical Society of America, 123, 27332749. 10.1121/1.2832337.
  41. Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., and Bright, J. (2020). SciPy 1.0: Fundamental algorithms for scientific computing in Python. Nature Methods, 17(3), 261272. 10.1038/s41592-020-0772-5.
  42. Wallmark, Z. (2018). The sound of evil: Timbre, body, and sacred violence in death metal. In R. Fink, M. Latour, and Z. Wallmark (Eds.), The Relentless Pursuit of Tone: Timbre in Popular Music (pp. 6587). Oxford University Press. 10.1093/oso/9780199985227.003.0004.
  43. Xu, Y., Wang, W., Cui, H., Xu, M., and Li, M. (2022). Paralinguistic singing attribute recognition using supervised machine learning for describing the classical tenor solo singing voice in vocal pedagogy. EURASIP Journal on Audio, Speech, and Music Processing, 2022(1), 8. 10.1186/s13636-022-00240-z.
DOI: https://doi.org/10.5334/tismir.310 | Journal eISSN: 2514-3298
Language: English
Page range: 309 - 328
Submitted on: Jun 30, 2025
Accepted on: May 13, 2026
Published on: Jul 16, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Xuhong Qiu, Emilia Parada‑Cabaleiro, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.