Skip to main content
Have a personal or library account? Click to login
Data‑Driven Analysis of Musical Form and Harmonic Structure in AI‑Generated Popular Music: A Case Study with Suno and Udio Cover

Data‑Driven Analysis of Musical Form and Harmonic Structure in AI‑Generated Popular Music: A Case Study with Suno and Udio

Open Access
|Sep 2026

References

  1. Agostinelli, A., Denk, T. I., Borsos, Z., Engel, J., Verzetti, M., Caillon, A., Huang, Q., Jansen, A., Roberts, A., Tagliasacchi, M., Sharifi, M., Zeghidour, N., and Frank, C. (2023). Musiclm: Generating music from text. arXiv preprint arXiv:2301.11325. 10.48550/arXiv.2301.11325.
  2. Ariza, C. (2009). The interrogator as critic: The Turing test and the evaluation of generative music systems. Computer Music Journal, 33(2), 4870. 10.1162/comj.2009.33.2.48.
  3. Benjamini, Y., and Hochberg, Y. (1995). Controlling the false discovery rate: A practical and powerful approach to multiple testing. Journal of the Royal Statistical Society: Series B (Methodological), 57(1), 289300. 10.1111/j.2517-6161.1995.tb02031.x.
  4. Bertin‑Mahieux, T., Ellis, D. P., Whitman, B., and Lamere, P. (2011). The million song dataset. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), pp. 591596. 10.5281/zenodo.1415820.
  5. Biamonte, N. (2010). Triadic modal and pentatonic patterns in rock music. Music Theory Spectrum, 32(2), 95110. 10.1525/mts.2010.32.2.95.
  6. Böck, S., Korzeniowski, F., Schlüter, J., Krebs, F., and Widmer, G. (2016). Madmom: A new Python audio and music signal processing library. In Proceedings of the 24th ACM International Conference on Multimedia (ACM‑MM), pp. 11741178. ACM. 10.1145/2964284.2973795.
  7. Bommasani, R., Creel, K. A., Kumar, A., Jurafsky, D., and Liang, P. S. (2022). Picking on the same person: Does algorithmic monoculture lead to outcome homogenization? Advances in Neural Information Processing Systems, 35, 36633678. 10.52202/068431-0265.
  8. Buisson, M., McFee, B., and Essid, S. (2024). Using pairwise link prediction and graph attention networks for music structure analysis. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), pp. 207214. 10.5281/zenodo.14877311.
  9. Buisson, M., McFee, B., Essid, S., and Crayencour, H.‑C. (2022). Learning multi‑level representations for hierarchical music structure analysis. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), pp. 591597. 10.5281/zenodo.7343060.
  10. Burgoyne, J. A., Wild, J., and Fujinaga, I. (2011). An expert ground truth set for audio chord recognition and music analysis. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), Vol. 11, pp. 633638. 10.5281/zenodo.1417547.
  11. Casini, L., Cros Vila, L., Dalmazzo, D., Kaila, A.‑K., and Sturm, B. L. T. (2026). Data‑driven analysis of text‑conditioning in AI‑generated music: A case study with Suno and Udio. Transactions of the International Society for Music Information Retrieval, 9(1), 194209. 10.5334/tismir.273.
  12. Collins, N. (2025). Recording artist career comparison through audio content analysis. Royal Society Open Science, 12(7), 241647. 10.1098/rsos.241647.
  13. Copet, J., Kreuk, F., Gat, I., Remez, T., Kant, D., Synnaeve, G., Adi, Y., and Défossez, A. (2023). Simple and controllable music generation. Advances in Neural Information Processing Systems, 36, 4770447720. 10.52202/075280-2066.
  14. Cros Vila, L., Sturm, B., Casini, L., and Dalmazzo, D. (2025). The AI music arms race: On the detection of AI‑generated music. Transactions of the International Society for Music Information Retrieval, 8(1), 179194. 10.5334/tismir.254.
  15. Cuthbert, M. S., and Ariza, C. (2010). music21: A toolkit for computer‑aided musicology and symbolic music data. In Proceedings of the 11th International Society for Music Information Retrieval Conference (ISMIR), pp. 637642. 10.5281/zenodo.1416114.
  16. Déguernel, K., and Sturm, B. L. T. (2023). Bias in favour or against computational creativity: A survey and reflection on the importance of socio‑cultural context in its evaluation. In Proceedings of the International Conference on Computational Creativity (ICCC).
  17. Dervakos, E., Filandrianos, G., and Stamou, G. (2021). Heuristics for evaluation of AI generated music. In 2020 25th International Conference on Pattern Recognition (ICPR), pp. 91649171. IEEE. 10.1109/ICPR48806.2021.9413310.
  18. European Union. (2019, May 17). Directive (EU) 2019/790 of the European Parliament and of the Council of 17 April 2019 on copyright and related rights in the Digital Single Market. Official Journal of the European Union, L 130, 92125. http://data.europa.eu/eli/dir/2019/790/oj.
  19. Evans, Z., Parker, J. D., Carr, C., Zukowski, Z., Taylor, J., and Pons, J. (2025). Stable audio open. In ICASSP 2025–2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 15. IEEE. 10.1109/ICASSP49660.2025.10888461.
  20. Foote, J. (2000). Automatic audio segmentation using a measure of audio novelty. In 2000 IEEE International Conference on Multimedia and Expo (ICME 2000) Proceedings. Latest Advances in the Fast Changing World of Multimedia (Cat. No.00TH8532), Vol. 1, pp. 452455. IEEE. 10.1109/ICME.2000.869637.
  21. Humphrey, E. J., and Bello, J. P. (2015). Four timely insights on automatic chord estimation. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), Vol. 10, pp. 673679. 10.5281/zenodo.1417549.
  22. Jiang, J., Chen, K., Li, W., and Xia, G. (2019). Large‑vocabulary chord transcription via chord structure decomposition. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), pp. 644651. 10.5281/zenodo.3527892.
  23. Kim, T., and Nam, J. (2023). All‑in‑one metrical and functional structure analysis with neighborhood attentions on demixed audio. In IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA). 10.1109/WASPAA58266.2023.10248148.
  24. Levy, M., and Sandler, M. (2008). Structural segmentation of musical audio by constrained clustering. IEEE Transactions on Audio, Speech, and Language Processing, 16(2), 318326. 10.1109/TASL.2007.910781.
  25. Lin, Z., Ehsan, U., Agarwal, R., Dani, S., Vashishth, V., and Riedl, M. (2023). Beyond prompts: Exploring the design space of mixed‑initiative co‑creativity systems. In Proceedings of the 14th International Conference on Computational Creativity (ICCC), pp. 6473. https://computationalcreativity.net/iccc23/.
  26. Marmoret, A., Cohen, J. E., and Bimbot, F. (2023). Barwise music structure analysis with the correlation block‑matching segmentation algorithm. Transactions of the International Society for Music Information Retrieval, 6(1), 167185. 10.5334/tismir.167.
  27. Mauch, M., and Dixon, S. (2009). Simultaneous estimation of chords and musical context from audio. IEEE Transactions on Audio, Speech, and Language Processing, 18(6), 12801289. 10.1109/TASL.2009.2032947.
  28. Mauch, M., and Dixon, S. (2010). Approximate note transcription for the improved identification of difficult chords. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), pp. 135140. 10.5281/zenodo.1416598.
  29. Mauch, M., MacCallum, R. M., Levy, M., and Leroi, A. M. (2015). The evolution of popular music: USA 1960–2010. Royal Society Open Science, 2(5), 150081. 10.1098/rsos.150081.
  30. McFee, B., and Bello, J. P. (2017). Structured training for large‑vocabulary chord recognition. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), pp. 188194. 10.5281/zenodo.1414880.
  31. McFee, B., and Ellis, D. P. (2014). Learning to segment songs with ordinal linear discriminant analysis. In 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 51975201. IEEE. 10.1109/ICASSP.2014.6854594.
  32. McFee, B., Raffel, C., Liang, D., Ellis, D. P., McVicar, M., Battenberg, E., and Nieto, O. (2025). librosa. 10.5281/zenodo.15006942.
  33. Müller, M. (2021). Fundamentals of Music Processing: Using Python and Jupyter Notebooks. Springer. 10.1007/978-3-030-69808-9.
  34. Ni, Y., McVicar, M., Santos‑Rodriguez, R., and De Bie, T. (2012). An end‑to‑end machine learning system for harmonic analysis of music. IEEE Transactions on Audio, Speech, and Language Processing, 20(6), 17711783. 10.1109/TASL.2012.2188516.
  35. Nieto, O., Mysore, G. J., Wang, C.‑I., Smith, J. B., Schlüter, J., Grill, T., and McFee, B. (2020). Audio‑based music structure analysis: Current trends, open challenges, and applications. Transactions of the International Society for Music Information Retrieval, 3(1). 10.5334/tismir.54.
  36. Odekerken, D., Koops, H. V., and Volk, A. (2021). Improving audio chord estimation by alignment and integration of crowd‑sourced symbolic music. Transactions of the International Society for Music Information Retrieval, 4(1), 141155. 10.5334/tismir.81.
  37. Oramas, S., Barbieri, F., Nieto, O., and Serrà, X. (2018). Multimodal deep learning for music genre classification. Transactions of the International Society for Music Information Retrieval, 1(1), 422. 10.5334/tismir.10.
  38. Park, J., Choi, K., Jeon, S., Kim, D., and Park, J. (2019). A bi‑directional transformer for musical chord recognition. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), pp. 620627. 10.5281/zenodo.3527886.
  39. Pasquier, P., Burnett, A., Thomas, N. G., Maxwell, J. B., Eigenfeldt, A., and Loughin, T. (2016). Investigating listener bias against musical metacreativity. In Proceedings of the International Conference on Computational Creativity (ICCC).
  40. Pauwels, J., and Peeters, G. (2013). Evaluating automatically estimated chord sequences. In 2013 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 749753. IEEE. 10.1109/ICASSP.2013.6637748.
  41. Peeters, G. (2023). Self‑similarity‑based and novelty‑based loss for music structure analysis. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), pp. 749756. 10.5281/zenodo.10265397.
  42. Pelly, L. (2025). Mood Machine: The Rise of Spotify and the Costs of the Perfect Playlist. Simon and Schuster.
  43. Poltronieri, A., Serrà, X., and Rocamora, M. (2025). From discord to harmony: Decomposed consonance‑based training for improved audio chord estimation. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), pp. 492500. 10.5281/zenodo.17811559.
  44. Poppe, T., Lopes, L., and Figueiredo, F. (2025). The polyphonic audio to roman corpus. In Proceedings of the 12th International Conference on Digital Libraries for Musicology, pp. 7280. 10.1145/3748336.3748345.
  45. Ragot, M., Martin, N., and Cojean, S. (2020). AI‑generated vs. human artworks: A perception bias towards artificial intelligence? In Proceedings of the Conference on Human Factors in Computing Systems (CHI). 10.1145/3334480.3382892.
  46. Richards, M. (2017). Tonal ambiguity in popular music’s axis progressions. Music Theory Online, 23(3). 10.30535/mto.23.3.6.
  47. Rouard, S., Massa, F., and Défossez, A. (2023). Hybrid transformers for music source separation. In ICASSP 2023–2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 15. IEEE. 10.1109/ICASSP49357.2023.10096956.
  48. Salamon, J., Nieto, O., and Bryan, N. J. (2021). Deep embeddings and section fusion improve music segmentation. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), pp. 594601. 10.5281/zenodo.5624371.
  49. Serrà, J., Corral, Á., Boguñá, M., Haro, M., and Arcos, J. L. (2012a). Measuring the evolution of contemporary western popular music. Scientific Reports, 2(1), 16. 10.1038/srep00521.
  50. Serrà, J., Müller, M., Grosche, P., and Arcos, J. L. (2012b). Unsupervised detection of music boundaries by time series structure features. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 26, pp. 16131619. 10.1609/aaai.v26i1.8328.
  51. Shank, D. B., Stefanik, C., Stuhlsatz, C., Kacirek, K., and Belfi, A. M. (2022). AI composer bias: Listeners like music less when they think it was composed by an AI. Journal of Experimental Psychology: Applied. 10.1037/xap0000447.
  52. Sturm, B. L., and Ben‑Tal, O. (2017). Taking the models back to music practice: Evaluating generative transcription models built using deep learning. Journal of Creative Music Systems, 2(1). 10.5920/jcms.2017.09.
  53. Sturm, B. L. T., Déguernel, K., Huang, R. S., Kaila, A.‑K., Jääskeläinen, P., Kanhov, E., Cros Vila, L., Dalmazzo, D., Casini, L., Bown, O. R., Collins, N., Drott, E., Sterne, J., Holzapfel, A., and Ben‑Tal, O. (2024). AI music studies: Preparing for the coming flood. In AIMC 2024, Oxford, United Kingdom. 10.5281/zenodo.15110181.
  54. Tan, S. (2024). Are we all musicians now? Authenticity, musicianship, and AI music generator Suno. 10.31235/osf.io/4nt8z.
  55. Temperley, D., and Clercq, T. d. (2013). Statistical analysis of harmony and melody in rock music. Journal of New Music Research, 42(3), 187204. 10.1080/09298215.2013.788039.
  56. Wang, J.‑C., Hung, Y.‑N., and Smith, J. B. (2022). To catch a chorus, verse, intro, or anything else: Analyzing a song with structural functions. In ICASSP 2022–2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 416420. IEEE. 10.1109/ICASSP43922.2022.9747252.
  57. Waseem Akram, M., Dettori, S., Colla, V., and Carlo Buttazzo, G. (2025). Chordformer: A conformer‑based architecture for large‑vocabulary audio chord recognition. IEEE Transactions on Audio, Speech, and Language Processing. 10.1109/TASLPRO.2025.3646468.
  58. Wiggins, J. (2007). Compositional process in music. In International Handbook of Research in Arts Education, pp. 453476. Springer. 10.1007/978-1-4020-3052-9_29.
  59. Wiggins, J. (2016). Musical agency. In G. E. McPherson (Ed.), The Child as Musician: A Handbook of Musical Development, pp. 102121. Oxford University Press. 10.1093/acprof:oso/9780198744443.003.0006.
  60. Wu, S.‑L., and Yang, Y.‑H. (2020). The jazz transformer on the front line: Exploring the shortcomings of AI‑composed music through quantitative measures. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), pp. 142149. 10.5281/zenodo.4245390.
  61. Zlatkov, D., Ens, J., and Pasquier, P. (2023). Searching for human bias against AI‑composed music. In Proceedings of the International Conference on Artificial Intelligence in Music, Sound, Art and Design (EvoMUSART). 10.1007/978-3-031-29956-8_20.
DOI: https://doi.org/10.5334/tismir.348 | Journal eISSN: 2514-3298
Language: English
Page range: 491 - 509
Submitted on: Oct 8, 2025
Accepted on: Aug 6, 2026
Published on: Sep 10, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 David Dalmazzo, Laura Cros Vila, Luca Casini, Bob L.T. Sturm, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.