Skip to main content
Have a personal or library account? Click to login
Cross-Cultural Music Similarity: Bridging Human Perception, Signal Processing, and Foundation Models Cover

Cross-Cultural Music Similarity: Bridging Human Perception, Signal Processing, and Foundation Models

Open Access
|Jul 2026

References

  1. Baevski, A., Zhou, Y., Mohamed, A., and Auli, M. (2020). wav2vec 2.0: A framework for self‑supervised learning of speech representations. In Advances in Neural Information Processing Systems (NeurIPS), Vancouver, Canada.
  2. Bello, J. P. (2011). Measuring structural similarity in music. IEEE Transactions on Audio, Speech, and Language Processing, 19(7), 20132025. 10.1109/tasl.2011.2108287.
  3. Bello, J. P., Daudet, L., Abdallah, S., Duxbury, C., Davies, M., and Sandler, M. B. (2005). A tutorial on onset detection in music signals. IEEE Transactions on Speech and Audio Processing, 13(5), 10351047. 10.1109/tsa.2005.851998.
  4. Bello, J. P., Duxbury, C., Davies, M. E., and Sandler, M. B. (2004). On the use of phase and energy for musical onset detection in the complex domain. IEEE Signal Processing Letters, 11(6), 553556. 10.1109/lsp.2004.827951.
  5. Berenzweig, A., Logan, B., Ellis, D. P., and Whitman, B. (2004). A large‑scale evaluation of acoustic and subjective music‑similarity measures. Computer Music Journal, 28(2), 6376. 10.1162/014892604323112257.
  6. Castellon, R., Donahue, C., and Liang, P. (2021). Codified audio language modeling learns useful representations for music information retrieval. Proceedings of the 22nd International Society for Music Information Retrieval Conference (ISMIR), virtual (pp. 8896).
  7. Chu, Y., Xu, J., Yang, Q., Wei, X., Wei, X., Guo, Z., Leng, Y., Lv, Y., He, J., Lin, J., Zhou, C., and Zhou, J. (2024). Qwen2‑audio technical report. arXiv preprint arXiv:2407.10759.
  8. Chu, Y., Xu, J., Zhou, X., Yang, Q., Zhang, S., Yan, Z., Zhou, C., and Zhou, J. (2023). Qwen‑audio: Advancing universal audio understanding via unified large‑scale audio‑language models. arXiv preprint arXiv:2311.07919.
  9. Defferrard, M., Benzi, K., Vandergheynst, P., and Bresson, X. (2017). FMA: A dataset for music analysis. Proceedings of the 18th International Society for Music Information Retrieval Conference (ISMIR), Suzhou, China (pp. 316323).
  10. Devlin, J., Chang, M. W., Lee, K., and Toutanova, K. (2019). BERT: Pre‑training of deep bidirectional transformers for language understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL‑HLT), Minneapolis, USA (pp. 41714186).
  11. Dubnov, S. (2004). Generalization of spectral flatness measure for non‑gaussian linear processes. IEEE Signal Processing Letters, 11(8), 698701. 10.1109/lsp.2004.831663.
  12. Dzhambazov, G. B., Srinivasamurthy, A., Şentürk, S., and Serra, X. (2016). On the use of note onsets for improved lyrics‑to‑audio alignment in Turkish Makam music. Proceedings of the 17th International Society for Music Information Retrieval Conference (ISMIR), New York City, USA (pp. 716722).
  13. Eerola, T., Järvinen, T., Louhivuori, J., and Toiviainen, P. (2001). Statistical features and perceived similarity of folk melodies. Music Perception, 18(3), 275296. 10.1525/mp.2001.18.3.275.
  14. Flexer, A. (2014). On inter‑rater agreement in audio music similarity. Proceedings of the 15th International Society for Music Information Retrieval Conference (ISMIR 2014), Taipei, Taiwan (pp. 245250).
  15. Fouloulis, T., Pikrakis, A., and Cambouropoulos, E. (2013). Traditional asymmetric rhythms: A refined model of meter induction based on asymmetric meter templates. Proceedings of the Third International Workshop on Folk Music Analysis, Amsterdam, Netherlands (pp. 2832).
  16. García‑Benito, R. (2025). Beyond universality: Cultural diversity in music and its implications for sound design and sonification. arXiv preprint arXiv:2506.14877.
  17. Grosche, P., Müller, M., and Kurth, F. (2010). Cyclic tempogram—a mid‑level tempo representation for music signals. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Dallas, USA (pp. 55225525).
  18. Gómez, E., Herrera, P., and Gómez‑Martin, F. (2013). Computational Ethnomusicology: Perspectives and challenges. Journal of New Music Research, 42(2), 111112. 10.1080/09298215.2013.818038.
  19. Harte, C., Sandler, M., and Gasser, M. (2006). Detecting harmonic change in musical audio. Proceedings of the 1st ACM Workshop on Audio and Music Computing Multimedia, Santa Barbara, USA (pp. 2126).
  20. Hawthorne, C., Stasyuk, A., Roberts, A., Simon, I., Huang, C.‑Z. A., Dieleman, S., Elsen, E., Engel, J., and Eck, D. (2018). Enabling factorized piano music modeling and generation with the MAESTRO dataset. arXiv preprint arXiv:1810. 12247.
  21. Holzapfel, A., Sturm, B. L., and Coeckelbergh, M. (2018). Ethical dimensions of music information retrieval technology. Transactions of the International Society for Music Information Retrieval (TISMIR), 1(1), 4455. 10.5334/tismir.13.
  22. Huang, P.‑Y., Xu, H., Li, J., Baevski, A., Auli, M., Galuba, W., Metze, F., and Feichtenhofer, C. (2022). Masked autoencoders that listen. In Advances in Neural Information Processing Systems (NeurIPS), New Orleans, USA (pp. 2870828720).
  23. Järvelin, K., and Kekäläinen, J. (2002). Cumulated gain‑based evaluation of IR techniques. ACM Transactions on Information Systems, 20(4), 422446. 10.1145/582415.582418.
  24. Kanatas, A.‑N., Papaioannou, C., and Potamianos, A. (2025). CultureMERT: Continual pre‑training for cross‑cultural music representation learning. Proceedings of the 26th International Society for Music Information Retrieval Conference (ISMIR), Daejeon, South Korea.
  25. Karaosmanoğlu, M. K., Bozkurt, B., Holzapfel, A., and Serra, X. (2014). A symbolic dataset of Turkish Makam music phrases. Proceedings of the 4th International Workshop on Folk Music Analysis (FMA), Istanbul, Turkey (pp. 1014).
  26. Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., and Liu, T.‑Y. (2017). LightGBM: A highly efficient gradient boosting decision tree. In Advances in Neural Information Processing Systems (NeurIPS), Long Beach, USA (pp. 31463154).
  27. Kendall, M. G. (1938). A new measure of rank correlation. Biometrika, 30(1–2), 8193. 10.1093/biomet/30.1-2.81.
  28. Klapuri, A., and Davy, M (Eds.). (2006). Signal Processing Methods for Music Transcription. Springer Verlag.
  29. Knees, P., and Schedl, M. (2013). A survey of music similarity and recommendation from music context data. ACM Transactions on Multimedia Computing, Communications, and Applications, 10(1), 121. 10.1145/2542205.2542206.
  30. Kouki, P., Fakhraei, S., Foulds, J., Eirinaki, M., and Getoor, L. (2015). HyPER: A flexible and extensible probabilistic framework for hybrid recommender systems. Proceedings of the 9th ACM Conference on Recommender Systems (RecSys), Vienna, Austria (pp. 99106).
  31. Kroher, N., Díaz‑Báñez, J.‑M., Mora, J., and Gómez, E. (2016). Corpus COFLA: A research corpus for the computational study of flamenco music. ACM Journal on Computing and Cultural Heritage, 9(2), 121. 10.1145/2875428.
  32. Krumhansl, C. L., and Castellano, M. A. (1983). Dynamic processes in music perception. Memory & Cognition, 11(4), 325334. 10.3758/bf03202445.
  33. Krumhansl, C. L., and Kessler, E. J. (1982). Tracing the dynamic changes in perceived tonal organization in a spatial representation of musical keys. Psychological Review, 89(4), 334343. 10.1037/0033-295x.89.4.334.
  34. Kruskal, J. B. (1964). Nonmetric multidimensional scaling: A numerical method. Psychometrika, 29(2), 115129. 10.1007/bf02289694.
  35. Lamont, A., and Dibben, N. (2001). Motivic structure and the perception of similarity. Music Perception, 18(3), 245274. 10.1525/mp.2001.18.3.245.
  36. Lartillot, O., and Toiviainen, P. (2007). A matlab toolbox for musical feature extraction from audio. Proceedings of the 10th International Conference on Digital Audio Effects (DAFx), Bordeaux, France (p. 244).
  37. Law, E., West, K., Mandel, M., Bay, M., and Downie, J. S. (2009). Evaluation of algorithms using games: The case of music tagging. Proceedings of the 10th International Society for Music Information Retrieval Conference (ISMIR), Kobe, Japan (pp. 387392).
  38. Lewin, D. (1982). Transformational techniques in atonal and other music theories. Perspectives of New Music, 21(1/2), 312371. 10.2307/832879.
  39. Li, Y., Yuan, R., Zhang, G., Ma, Y., Chen, X., Yin, H., Xiao, C., Lin, C., Ragni, A., Benetos, E., Gyenge, N., Dannenberg, R., Liu, R., Chen, W., Xia, G., Shi, Y., Huang, W., Wang, Z., Guo, Y., and Fu, J. (2024). MERT: Acoustic music understanding model with large‑scale self‑supervised training. International Conference on Learning Representations (ICLR).
  40. Logan, B. (2000). Mel frequency cepstral coefficients for music modeling. Proceedings of the 1st International Symposium on Music Information Retrieval (ISMIR), Plymouth, USA.
  41. Ma, Y., Øland, A., Ragni, A., Del Sette, B. M., Saitis, C., Donahue, C., Lin, C., Plachouras, C., Benetos, E., Shatri, E., Morreale, F., Zhang, G., Fazekas, G., Xia, G., Zhang, H., Manco, I., Huang, J., Guinot, J., Lin, L., and Wang, Z, … (2024). Foundation models for music: A survey. arXiv preprint arXiv:2408.14340.
  42. Mauch, M., and Dixon, S. (2014). PYIN: A fundamental frequency estimator using probabilistic threshold distributions. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Florence, Italy (pp. 659663).
  43. McDermott, H. J. (2004). Music perception with cochlear implants: A review. Trends in Amplification, 8(2), 4982. 10.1177/108471380400800203.
  44. McFee, B., Raffel, C., Liang, D., Ellis, D. P., McVicar, M., Battenberg, E., and Nieto, O. (2015). librosa: Audio and music signal analysis in Python. Proceedings of the 14th Python in Science Conference (SciPy), Austin, USA (pp. 1824).
  45. Mehr, S. A., Singh, M., Knox, D., Ketter, D., Pickens‑Jones, D., Atwood, S., Lucas, C., Jacoby, N., Egner, A., Hopkins, E. J., Howard, R. M., Hartshorne, J. K., Jennings, M. V., Simson, J., Bainbridge, C. M., Pinker, S., O’Donnell, T. J., Krasnow, M. M., and Glowacki, L. (2019). Universality and diversity in human song. Science, 366(6468), eaax0868. 10.1126/science.aax0868.
  46. Mehta, A., Chauhan, S., Djanibekov, A., Kulkarni, A., Xia, G., and Choudhury, M. (2025). Music for all: Representational bias and cross‑cultural adaptability of music generation models. In Findings of the Association for Computational Linguistics (NAACL), Albuquerque, USA (pp. 45694585).
  47. Meredith, D. (2014). Using point‑set compression to classify folk songs. Proceedings of the 4th International Workshop on Folk Music Analysis (FMA), Istanbul, Turkey (pp. 2935).
  48. Müller, M., Arzt, A., Balke, S., Dorfer, M., and Widmer, G. (2018). Cross‑modal music retrieval and applications: An overview of key methodologies. IEEE Signal Processing Magazine, 36(1), 5262. 10.1109/msp.2018.2868887.
  49. Müller, M., and Ewert, S. (2011). Chroma Toolbox: MATLAB implementations for extracting variants of chroma‑based audio features. Proceedings of the 12th International Society for Music Information Retrieval Conference (ISMIR), Miami, USA (pp. 215220).
  50. Nawrot, E. S. (2003). The perception of emotional expression in music: Evidence from infants, children and adults. Psychology of Music, 31(1), 7592. 10.1177/0305735603031001325.
  51. Ozaki, Y., Tierney, A., Pfordresher, P. Q., McBride, J. M., Benetos, E., Proutskova, E., Chiba, G., Liu, F., Jacoby, N., Purdy, S. C., Opondo, P., Fitch, W. T., Hegde, S., Rocamora, M., Thorne, R., Nweke, F., Sadaphal, D. P., Sadaphal, P. M., Hadavi, S., and Savage, P. E, … (2024). Globally, songs and instrumental melodies are slower and higher and use more stable pitches than speech: A registered report. Science Advances, 10(20), eadm9797. 10.1126/sciadv.adm9797.
  52. Panteli, M. (2018). Computational Analysis of World Music Corpora. (Doctoral dissertation). Queen Mary University of London.
  53. Panteli, M., Benetos, E., and Dixon, S. (2018). A review of manual and computational approaches for the study of world music corpora. Journal of New Music Research, 47(2), 176189. 10.1080/09298215.2017.1418896.
  54. Papaioannou, C., Benetos, E., and Potamianos, A. (2025). Universal music representations? Evaluating foundation models on world music corpora. Proceedings of the 26th International Society for Music Information Retrieval Conference (ISMIR), Daejeon, South Korea.
  55. Papaioannou, C., Valiantzas, I., Giannakopoulos, T., Kaliakatsos‑Papakostas, M., and Potamianos, A. (2022). A dataset for Greek traditional and folk music: Lyra. Proceedings of the 23rd International Society for Music Information Retrieval Conference (ISMIR), Bengaluru, India (pp. 377383).
  56. Patel, A. D. (2010). Music, Language, and the Brain. Oxford University Press.
  57. Plachouras, C. (2023). Beyond Benchmarks: A Toolkit for Music Audio Representation Evaluation. (Master’s thesis). Universitat Pompeu Fabra.
  58. Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. (2021). Learning transferable visual models from natural language supervision. Proceedings of the 38th International Conference on Machine Learning (ICML) virtual, 139, 87488763. proceedings.mlr.press/v139/radford21a.html.
  59. Radinsky, K., and Ailon, N. (2011). Ranking from pairs and triplets: Information quality, evaluation methods and query complexity. Proceedings of the 4th ACM International Conference on Web Search and Data Mining (WSDM), Hong Kong, China (pp. 105114).
  60. Repetto, R. C., Pretto, N., Chaachoo, A., Bozkurt, B., and Serra, X. (2018). An open corpus for the computational research of Arab‑Andalusian music. Proceedings of the 5th International Workshop on Digital Libraries for Musicology (DLFM), Paris, France (pp. 7886).
  61. Repetto, R. C., and Serra, X. (2014). Creating a corpus of jingju (Beijing opera) music and possibilities for melodic analysis. Proceedings of the 15th International Society for Music Information Retrieval Conference (ISMIR), Taipei, Taiwan (pp. 313318).
  62. Savage, P. E., Brown, S., Sakai, E., and Currie, T. E. (2015). Statistical universals reveal the structures and functions of human music. Proceedings of the National Academy of Sciences, 112, 89878992. 10.1073/pnas.1414495112.
  63. Schedl, M., Gómez, E., and Urbano, J. (2014). Music information retrieval: Recent developments and applications. Foundations and Trends in Information Retrieval, 8(2–3), 127261. 10.1561/1500000042.
  64. Şentürk, S. (2016). Computational Analysis of Audio Recordings and Music Scores for the Description and Discovery of Ottoman‑Turkish Makam Music. (Doctoral dissertation). Universitat Pompeu Fabra.
  65. Serra, X. (2014). Creating research corpora for the computational study of music: The case of the CompMusic project. Proceedings of the AES 53rd International Conference on Semantic Audio, London, UK.
  66. Serra, X., Magas, M., Benetos, E., Chudy, M., Dixon, S., Flexer, A., Gómez, E., Gouyon, F., Herrera, P., Jordà, S., Paytuvi, O., Peeters, G., Schlüter, J., Vinet, H., and Widmer, G. (2013). Roadmap for Music Information Research. Music Technology Group, Universitat Pompeu Fabra.
  67. Siedenburg, K., Fujinaga, I., and McAdams, S. (2016). A comparison of approaches to timbre descriptors in music information retrieval and music psychology. Journal of New Music Research, 45(1), 2741. 10.1080/09298215.2015.1132737.
  68. Spearman, C. (1904). The proof and measurement of association between two things. The American Journal of Psychology, 15(1), 72101. 10.2307/1412159.
  69. Srinivasamurthy, A., Holzapfel, A., and Serra, X. (2014a). In search of automatic rhythm analysis methods for Turkish and Indian art music. Journal of New Music Research, 43(1), 94114. 10.1080/09298215.2013.879902.
  70. Srinivasamurthy, A., Koduri, G. K., Gulati, S., Ishwar, V., and Serra, X. (2014b). Corpora for music information research in Indian art music. Proceedings of the 40th International Computer Music Conference (ICMC), Athens, Greece.
  71. Sturm, B. L. (2013). Classification accuracy is not enough ‑ on the evaluation of music genre recognition systems. Journal of Intelligent Information Systems, 41(3), 371406. 10.1007/s10844-013-0250-y.
  72. Tian, H., Lattner, S., and Saitis, C. (2025). Assessing the alignment of audio representations with timbre similarity ratings. Proceedings of the 26th International Society for Music Information Retrieval Conference (ISMIR), Daejeon, South Korea.
  73. Typke, R. (2007). Music Retrieval Based on Melodic Similarity. (Doctoral dissertation). Utrecht University.
  74. Tzanetakis, G., and Cook, P. (2002). Musical genre classification of audio signals. IEEE Transactions on Speech and Audio Processing, 10(5), 293302. 10.1109/tsa.2002.800560.
  75. Uyar, B., Atli, H. S., Şentürk, S., Bozkurt, B., and Serra, X. (2014). A corpus for computational research of Turkish Makam music. Proceedings of the 4th International Workshop on Folk Music Analysis (FMA), Istanbul, Turkey (pp. 17).
  76. Volk, A., de Haas, W. B., and Van Kranenburg, P. (2012). Towards modelling variation in music as foundation for similarity. Proceedings of the 12th International Conference on Music Perception and Cognition (ICMPC) (pp. 10851094).
  77. Volk, A., and Van Kranenburg, P. (2012). Melodic similarity among folk songs: An annotation study on similarity‑based categorization in music. Musicae Scientiae, 16(3), 317339. 10.1177/1029864912448329.
  78. Weisberg, S. (2005). Applied Linear Regression (3rd ed., Vol. 528). John Wiley & Sons.
  79. Won, M., Ferraro, A., Bogdanov, D., and Serra, X. (2020). Evaluation of CNN‑based automatic music tagging models. CoRR. abs/2006.00751
  80. Wu, Y., Chen, K., Zhang, T., Hui, Y., Nezhurina, M., Berg‑Kirkpatrick, T., and Dubnov, S. (2023). Large‑scale contrastive language‑audio pretraining with feature fusion and keyword‑to‑caption augmentation. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece (pp. 15).
DOI: https://doi.org/10.5334/tismir.341 | Journal eISSN: 2514-3298
Language: English
Page range: 347 - 368
Submitted on: Sep 15, 2025
Accepted on: Mar 29, 2026
Published on: Jul 21, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Charilaos Papaioannou, Emmanouil Benetos, Alexandros Potamianos, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.