Skip to main content
Have a personal or library account? Click to login
Attention‑Enhanced Convolutional Neural Network for Music Genre Classification Cover

Attention‑Enhanced Convolutional Neural Network for Music Genre Classification

Open Access
|Aug 2026

References

  1. Ananthanarayana, R. M., Bhattacharjee, A., and Rao, P. (2023). Four‑way classification of tabla strokes with transfer learning using western drums. Transactions of the International Society for Music Information Retrieval, 6(1), 103116. 10.5334/tismir.150.
  2. Ashraf, M., Abid, F., Din, I. U., Rasheed, J., Yesiltepe, M., Yeo, S. F., and Ersoy, M. T. (2023). A hybrid CNN and rnn variant model for music classification. Applied Sciences, 13(3), 1476. 10.3390/app13031476.
  3. Asif, A., Majid, M., and Anwar, S. M. (2019). Human stress classification using EEG signals in response to music tracks. Computers in Biology and Medicine, 107, 182196. 10.1016/j.compbiomed.2019.02.015.
  4. Jayakanthan, A. P., Asswin, C. R., Kumar, Dharshan., S, K., Dora, A., Ravi, V., Sowmya, V., Gopalakrishnan, E. A., and Soman, K. P. (2023). Transfer learning approach for pediatric pneumonia diagnosis using channel attention deep cnn architectures. Engineering Applications of Artificial Intelligence, 123, 106416. 10.1016/j.engappai.2023.106416.
  5. Bahuleyan, H. (2018). Music genre classification using machine learning techniques. arXiv Preprint arXiv:1804.01149.
  6. Balaji, A. J., Harish Ram, D., and Nair, B. B. (2019). A deep learning approach to electric energy consumption modeling. Journal of Intelligent & Fuzzy Systems, 36(5), 40494055. 10.3233/jifs-169965.
  7. Berenzweig, A., Logan, B., Ellis, D. P. W., and Whitman, B. (2003). Anchor space for classification and similarity measurement of music. In Proceedings of the IEEE International Conference on Multimedia and Expo (ICME), Baltimore, MD, USA (pp. 2932). IEEE.
  8. Chen, J., Ma, X., Li, S., Ma, S., Zhang, Z., and Ma, X. (2024). A hybrid parallel computing architecture based on CNN and transformer for music genre classification. Electronics, 13(16), 3313. 10.3390/electronics13163313.
  9. Cheng, Y.‑H., Chang, P.‑C., Nguyen, D.‑M., and Kuo, C.‑N. (2020). Automatic music genre classification based on CRNN. Engineering Letters, 29(1).
  10. Cheng, Y.‑H., and Kuo, C.‑N. (2022). Machine learning for music genre classification using visual mel spectrum. Mathematics, 10(23), 4427. 10.3390/math10234427.
  11. da, Silva., M, A. C., Coelho, M. A. N., and Neto, R. F. (2020). A music classification model based on metric learning applied to mp3 audio files. Expert Systems with Applications, 144, 113071. 10.1016/j.eswa.2019.113071.
  12. Dhall, A., Srinivasa Murthy, Y., and Koolagudi, S. G. (2021). Music genre classification with convolutional neural networks and comparison with f, q, and mel spectrogram‑based images. In Advances in Speech and Music Technology (pp. 235248). Springer. 10.1007/978-981-33-6881-1_20">http://doi.org/10.1007/978-981-33-6881-1_20.
  13. Eerola, T., Vuoskoski, J. K., Peltola, H.‑R., Putkinen, V., and Schäfer, K. (2018). An integrative review of the enjoyment of sadness associated with music. Physics of Life Reviews, 25, 100121. 10.1016/j.plrev.2017.11.016.
  14. Ferraro, A., Bogdanov, D., and Serra, X. (2020). How low can you go? Reducing frequency and time resolution in current cnn architectures for music auto‑tagging. In Proceedings of the European Signal Processing Conference (EUSIPCO) (pp. 186190). IEEE.
  15. Flexer, A. (2007). A closer look on artist filters for musical genre classification. World, 19(122), 1617. https://ismir2007.ismir.net/proceedings/ISMIR2007_p341_flexer.pdf.
  16. Folorunso, S. O., Afolabi, S. A., and Owodeyi, A. B. (2022). Dissecting the genre of Nigerian music with machine learning models. Journal of King Saud University – Computer and Information Sciences, 34(8), 62666279. 10.1016/j.jksuci.2021.07.009.
  17. Fulzele, P., Singh, R., Kaushik, N., and Pandey, K. (2018). A hybrid model for music genre classification using LSTM and SVM. In 2018 Eleventh International Conference on Contemporary Computing (IC3) (pp. 13). IEEE.
  18. Ghosal, D., and Kolekar, M. H. (2018). Music genre recognition using deep neural networks and transfer learning. In Interspeech (pp. 20872091). 10.21437/interspeech.2018-2045.
  19. Han, K., Wang, Y., Chen, H., Chen, X., Guo, J., Liu, Z., Tang, Y., Xiao, A., Xu, C., Xu, Y., Yang, Z., Zhang, Y., and Tao, D. (2022). A survey on vision transformer. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(1), 87110. 10.1109/tpami.2022.3152247.
  20. Hasib, K. M., Tanzim, A., Shin, J., Faruk, K. O., Al Mahmud, J., and Mridha, M. F. (2022). Bmnet‑5: A novel approach of neural network to classify the genre of Bengali music based on audio features. IEEE Access, 10, 108545108563. 10.1109/access.2022.3213818.
  21. Hassani, A., Walton, S., Shah, N., Abuduweili, A., Li, J., and Shi, H. (2021). Escaping the big data paradigm with compact transformers. arXiv Preprint arXiv:2104.05704.
  22. Hongdan, W., SalmiJamali, S., Zhengping, C., Qiaojuan, S., and Le, R. (2022). An intelligent music genre analysis using feature extraction and classification using deep learning techniques. Computers and Electrical Engineering, 100, 107978. 10.1016/j.compeleceng.2022.107978.
  23. Jerzak, A. (2025). An accidental benchmark: The history, contingent power, and lasting traces of the GTZAN dataset. Digital Society, 4(2), 34. 10.1007/s44206-025-00191-w.
  24. Li, J., Han, L., Li, X., Zhu, J., Yuan, B., and Gou, Z. (2022). An evaluation of deep neural network models for music classification using spectrograms. Multimedia Tools and Applications, 81(4), 46214647. 10.1007/s11042-020-10465-9.
  25. Li, T., and Ogihara, M. (2005). Music genre classification with taxonomy. In 2005 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP’05), 5, 197200. 10.1109/ICASSP.2005.1416274.
  26. Li, T., and Ogihara, M. (2006). Toward intelligent music information retrieval. IEEE Transactions on Multimedia, 8(3), 564574. 10.1109/tmm.2006.870730.
  27. Liu, C., Feng, L., Liu, G., Wang, H., and Liu, S. (2021). Bottom‑up broadcast neural network for music genre classification. Multimedia Tools and Applications, 80, 73137331. 10.1007/s11042-020-09643-6.
  28. Liuwanyue, S. (2024). Course genres classification of music e‑learning platform based on deep learning big data intelligent processing algorithm. Entertainment Computing, 50, 100704. 10.1016/j.entcom.2024.100704.
  29. Markov, K., and Matsui, T. (2014). Music genre and emotion recognition using gaussian processes. IEEE Access, 2, 688697. 10.1109/access.2014.2333095.
  30. McKay, C., and Fujinaga, I. (2006). Musical genre classification: Is it worth pursuing and how can it be improved? In Proceedings of the 7th International Society for Music Information Retrieval Conference (ISMIR) (pp. 101106).
  31. Müller, M., Arzt, A., Balke, S., Dorfer, M., and Widmer, G. (2018). Cross‑modal music retrieval and applications: An overview of key methodologies. IEEE Signal Processing Magazine, 36(1), 5262. 10.1109/msp.2018.2868887.
  32. Nanni, L., Costa, Y. M., Lumini, A., Kim, M. Y., and Baek, S. R. (2016). Combining visual and acoustic features for music genre classification. Expert Systems with Applications, 45, 108117. 10.1016/j.eswa.2015.09.018.
  33. Nunes, I., Santana, M. A., Charron, N., Souza e, Silva., Simoes, H., Lins, C., Souza Sampaio, C., de., B, A., de Melo, A. M. N., da, Silva., V, T. C., de Lima, C. T., Córdula, N., Sarmento, A., Gomes, J. C., Moreno, G. M. M., Gusmão, C., and Dos Santos, W. P. (2024). Automatic identification of preferred music genres: An exploratory machine learning approach to support personalized music therapy. Multimedia Tools and Applications, 83(35), 8251582531. 10.1007/s11042-024-18826-4.
  34. Oramas, S., Barbieri, F., Nieto, O., and Serra, X. (2018). Multimodal deep learning for music genre classification. Transactions of the International Society for Music Information Retrieval, 1(1), 421. 10.5334/tismir.10.
  35. Panagakis, Y., Kotropoulos, C. L., and Arce, G. R. (2014). Music genre classification via joint sparse low‑rank representation of audio features. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 22(12), 19051917. 10.1109/taslp.2014.2355774.
  36. Pourmoazemi, N., and Maleki, S. (2024). A music recommender system based on compact convolutional transformers. Expert Systems with Applications, 255, 124473. 10.1016/j.eswa.2024.124473.
  37. Prabhakar, S. K., and Lee, S.‑W. (2023). Holistic approaches to music genre classification using efficient transfer and deep learning techniques. Expert Systems with Applications, 211, 118636. 10.1016/j.eswa.2022.118636.
  38. Pushparajan, M., Sreekumar, K., Ramachandran, K., and Kumar, C. S. (2022). Performance enhancement of raga classification systems using recursive feature elimination. In Third International Conference on Sustainable Computing, Singapore (pp. 533541). Springer Nature Singapore.
  39. Pushparajan, M., Sreekumar, K., Ramachandran, K., and Kumar, C. S. (2024). Data augmentation for improving the performance of raga (music genre) classification systems. In 5th International Conference on Electronics and Sustainable Communication Systems (ICESC) (pp. 14071412). IEEE.
  40. Ramírez, J., and Flores, M. J. (2020). Machine learning for music genre: Multifaceted review and experimentation with audioset. Journal of Intelligent Information Systems, 55(3), 469499. 10.1007/s10844-019-00582-9.
  41. Schneider, D., Korfhage, N., Mühling, M., Lüttig, P., and Freisleben, B. (2021). Automatic transcription of organ tablature music notation with deep neural networks. Transactions of the International Society for Music Information Retrieval, 4(1), 1428. 10.5334/tismir.77.
  42. Shao, X., Xu, C., and Kankanhalli, M. S. (2004). Unsupervised classification of music genre using hidden Markov model. In 2004 IEEE International Conference on Multimedia and Expo (ICME), 3, 20232026.
  43. Sturm, B. L. (2013). The GTZAN dataset: Its contents, its faults, their effects on evaluation, and its future use. arXiv Preprint arXiv:1306.1461.
  44. Sturm, B. L. (2014). The state of the art ten years after a state of the art: Future research in music information retrieval. Journal of New Music Research, 43(2), 147172. 10.1080/09298215.2014.894533.
  45. Tzanetakis, G., and Cook, P. (2002). Musical genre classification of audio signals. IEEE Transactions on Speech and Audio Processing, 10(5), 293302. 10.1109/tsa.2002.800560.
  46. Vatolkin, I., and McKay, C. (2022). Multi‑objective investigation of six feature source types for multimodal music classification. Transactions of the International Society for Music Information Retrieval, 5(1), 118. 10.5334/tismir.67.
  47. West, K., and Cox, S. (2005). Finding an optimal segmentation for audio genre classification. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR) (pp. 680685).
  48. Xu, C., Maddage, N. C., Shao, X., Cao, F., and Tian, Q. (2003). Musical genre classification using support vector machines. In In 2003 IEEE International Conference on Acoustics, Speech, and Signal Processing (Vol. 5, pp. V429). IEEE. 10.1109/icassp.2003.1199998.
  49. Yang, R., Feng, L., Wang, H., Yao, J., and Luo, S. (2020). Parallel recurrent convolutional neural networks‑based music genre classification method for mobile devices. IEEE Access, 8, 1962919637. 10.1109/access.2020.2968170.
  50. Yu, Y., Luo, S., Liu, S., Qiao, H., Liu, Y., and Feng, L. (2020). Deep attention based music genre classification. Neurocomputing, 372, 8491. 10.1016/j.neucom.2019.09.054.
  51. Zaman, K., Sah, M., Direkoglu, C., and Unoki, M. (2023). A survey of audio classification using deep learning. IEEE Access, 11, 106620106649. 10.1109/access.2023.3318015.
  52. Zhang, W., Lei, W., Xu, X., and Xing, X. (2016). Improved music genre classification with convolutional neural networks. In In Interspeech (pp. 33043308). 10.21437/interspeech.2016-1236.
  53. Zhao, H., Zhang, C., Zhu, B., Ma, Z., and Zhang, K. (2022). S3t: Self‑supervised pre‑training with swin transformer for music classification. In ICASSP 2022 – 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 606610). IEEE.
DOI: https://doi.org/10.5334/tismir.268 | Journal eISSN: 2514-3298
Language: English
Page range: 423 - 439
Submitted on: Apr 10, 2025
Accepted on: Jun 18, 2026
Published on: Aug 4, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 M Pushparajan, KT Sreekumar, KI Ramachandran, C Santhosh Kumar, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.