
Music Genre Classification with Multi-Modal Properties of Lyrics and Spectrograms
By: Janitha Madushan and Ruwan Weerasinghe
Abstract
Music Genre classification is widely used in online music streaming platforms. Deep learning has enabled extract-ing musical information more effectively, and there have been various research works done to improve their accuracy with power spectrogram images and lyrical features. This paper evaluated the optimum usage of multiple modalities such as lyrics and spectrogram images based on the richness of their features. Furthermore, it proposes a hybrid-fusion-based deep learning multi-modal, multi-class classifier, that employs Mel Spectrograms, Mel-Frequency Cepstral Coefficients, and Lyrics to classify musical genres more accurately. Finally, the proposed model benchmarked with 3 previous studies, with a prepossessed dataset from the Music4All dataset with country, jazz, metal, and pop genre classes and obtained the highest F1-Score of 0.72 for the proposed model.
DOI: https://doi.org/10.4038/icter.v18i2.7292 | Journal eISSN: 2550-2794
Language: English
Page range: 58 - 65
Published on: May 31, 2025
Published by: University of Colombo School of Computing
In partnership with: Paradigm Publishing Services
Keywords:
© 2025 Janitha Madushan, Ruwan Weerasinghe, published by University of Colombo School of Computing
This work is licensed under the Creative Commons Attribution 4.0 License.