
AI-Driven Emotional Analysis of Musical Structures for Enhanced Content Creation in the Modern Music Industry
Abstract
In our research, we experimented with identifying emotions in music tracks by utilizing a multimodal deep learning approach that incorporated lyrical and audio data. To obtain semantic representations of the lyrics, we used BERT embeddings. As for audio modality, mel-spectrograms were extracted and forwarded into a convolutional network. Both representations were projected to the latent space of a given size which allowed our model to process both lyrics and audio simultaneously and extract emotional features from textual and acoustic representations of songs. We compared multimodal solutions to models that operated on a single modality: audio or lyrics. Results of our experiments demonstrated that utilizing both modalities allows our network to converge faster and attain better overall performance for all categories present in our dataset. During our experiments, we consistently observed that there are patterns in music which express emotions that can be better captured when considered from both an audio and linguistic perspective. Applications of our research could allow for music recommendation, soundtrack-related solutions, and assistance in song writing/producing. There are several limitations to our study. For one, our dataset was quite limited in size. In addition, some categories were not as heavily represented as others. We also used generated text to represent song lyrics in some cases instead of using the actual lyrics of the songs. Future directions can include working with bigger datasets, better language understanding for music lyrics, and experimenting with different techniques for multimodal fusion.
© 2026 Valentin-Ionut-Cosmin DUMITRESCU, Gheorghe MILITARU, published by Bucharest University of Economic Studies
This work is licensed under the Creative Commons Attribution 4.0 License.