Table 1
Summary of musical datasets used in cross‑cultural similarity study.
| Dataset | Tradition | Hours | Recordings |
|---|---|---|---|
| MagnaTagATune (Law et al., 2009) | Western | 210 | 25,864 |
| FMA‑medium (Defferrard et al., 2017) | Western | 208 | 25,000 |
| CorpusCOFLA (Kroher et al., 2016) | Flamenco | 95 | 1,595 |
| Arab‑Andalusian (Repetto et al., 2018) | Spanish‑Arabic | 125 | 164 |
| Lyra (Papaioannou et al., 2022) | Greek | 80 | 1,570 |
| Turkish‑makam (Şentürk, 2016; Uyar et al., 2014) | Turkish | 359 | 5,297 |
| Hindustani (Srinivasamurthy et al., 2014b) | Indian | 343 | 1,204 |
| Carnatic (Srinivasamurthy et al., 2014b) | Indian | 500 | 2,612 |
| Jingju (Repetto and Serra, 2014) | Chinese | 71 | 864 |
Table 2
Summary Statistics of the Human Annotation Study. Comprehensive overview of participant demographics, study design parameters, and data collection metrics.
| Category | Count |
|---|---|
| Total participants | 125 |
| Unique annotated audio pairs | 1,130 |
| Unique annotated audio clips | 463 |
| Annotated audio pairs per participant | 10 |
| Annotated similarity types per audio pair | 3 |
| Participants unique countries of origin | 21 |
| Participants unique music training levels | 13 |
| Participants unique familiar music cultures | 58 |
Table 3
Signal‑Processing Features Summary. Overview of features extracted for each musical dimension in the cross‑cultural analysis framework.
| Dimension | Features |
|---|---|
| Melody | PYIN F0 extraction (Mauch and Dixon, 2014), dual‑resolution pitch classes (12/24 bins), melodic intervals, F0 statistics |
| Rhythm | Dynamic programming beat tracking (McFee et al., 2015), onset detection (Bello et al., 2004, 2005), inter‑onset/beat intervals, tempogram analysis (Grosche et al., 2010) |
| Harmony | CENS chroma features (Müller and Ewert, 2011) (12/24 bins), chord recognition, Krumhansl–Schmuckler key estimation (Krumhansl and Kessler, 1982), Tonnetz centroids (Harte et al., 2006) |
| Timbre | MFCCs (Logan, 2000; Tzanetakis and Cook, 2002) with delta features, spectral features (Klapuri and Davy, 2006), spectral flatness (Dubnov, 2004), statistical distributions |
Table 4
Foundation Models Summary. Overview of the seven foundation models used in cross‑cultural music similarity evaluation. CultureMERT‑TA employs task arithmetic to merge culture‑specific models in weight space.
| Model | Params | Key Characteristics |
|---|---|---|
| MERT‑95M (Li et al., 2024) | 95M | 12‑layer transformer, dual‑teacher masked acoustic modeling |
| MERT‑330M (Li et al., 2024) | 330M | 24‑layer transformer, expanded MERT variant |
| CultureMERT (Kanatas et al., 2025) | 95M | Continual pre‑training on Greek, Turkish, and Indian music |
| CultureMERT‑TA (Kanatas et al., 2025) | 95M | Task arithmetic cultural adaptation approach |
| CLAP‑Music (Wu et al., 2023) | 194M | Contrastive audio–text learning, music‑only training |
| CLAP‑Music&Speech (Wu et al., 2023) | 194M | Contrastive audio–text learning, music and speech data |
| Qwen2‑Audio (Chu et al., 2024) | 8.4B | Multimodal architecture with instruction tuning |

Figure 1
Cultural Similarity Matrix Across Datasets. Heat map visualization of human‑perceived cultural similarity ratings. Values represent mean cultural similarity ratings aggregated across all participant annotations for pairs between and within datasets. Clear cultural clusters emerge, with higher similarities (darker blue) indicating stronger cultural relationships.

Figure 2
Multidimensional Scaling Visualization of Musical Datasets. Two‑dimensional projection based on recommendation‑level similarity distances derived from human annotations, revealing clustering patterns across nine musical datasets. Dotted circles around each point represent internal diversity (inverse self‑similarity) within each dataset.
Table 5
Comprehensive Evaluation of Signal‑Processing Features and Foundation Models. Performance comparison against human similarity judgments across three similarity dimensions (overall musical, cultural, and recommendation‑level). Values are shown as percentages (%) for Triplet Agreement, NDCG, and MAE and as correlation values for Spearman and Kendall metrics. Arrows indicate whether higher () or lower () values represent better performance, with the best performance within each similarity dimension and metric shown in bold.
| Method | Triplet Agr. (%) | NDCG (%) | Spearman (−1, 1) | Kendall (−1, 1) | MAE (%) | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Similarity type | Overall | Cultural | Recomm. | Overall | Cultural | Recomm. | Overall | Cultural | Recomm. | Overall | Cultural | Recomm. | Overall | Cultural | Recomm. |
| Signal processing features | |||||||||||||||
| Melody | 61.5 | 61.1 | 60.7 | 88.4 | 87.6 | 86.8 | 0.15 | 0.14 | 0.15 | 0.14 | 0.12 | 0.13 | 29.5 | 30.5 | 30.9 |
| Rhythm | 51.3 | 52.1 | 50.3 | 85.8 | 84.0 | 84.0 | −0.00 | −0.01 | −0.02 | −0.00 | −0.01 | −0.02 | 32.5 | 34.3 | 34.6 |
| Harmony | 51.8 | 50.8 | 50.7 | 85.3 | 83.4 | 83.6 | 0.02 | −0.00 | 0.02 | 0.02 | 0.00 | 0.02 | 32.1 | 33.5 | 34.3 |
| Timbre | 54.2 | 54.7 | 55.6 | 86.1 | 84.8 | 85.3 | −0.03 | 0.04 | 0.04 | −0.03 | 0.03 | 0.03 | 35.2 | 36.4 | 36.3 |
| Foundation models | |||||||||||||||
| MERT‑95 | 59.8 | 59.7 | 60.0 | 88.2 | 87.1 | 87.3 | 0.06 | 0.09 | 0.10 | 0.05 | 0.08 | 0.08 | 31.3 | 32.4 | 32.3 |
| CultureMERT | 56.3 | 57.0 | 57.4 | 86.8 | 86.2 | 86.4 | 0.04 | 0.08 | 0.08 | 0.03 | 0.06 | 0.07 | 33.0 | 34.1 | 34.4 |
| CultureMERT‑TA | 55.1 | 55.8 | 56.5 | 86.6 | 86.0 | 86.4 | 0.02 | 0.06 | 0.06 | 0.01 | 0.05 | 0.05 | 33.6 | 34.6 | 34.8 |
| MERT‑330 | 57.6 | 57.3 | 58.7 | 87.8 | 86.5 | 86.9 | 0.08 | 0.05 | 0.09 | 0.06 | 0.04 | 0.08 | 35.0 | 35.6 | 35.7 |
| CLAP‑Music | 55.6 | 56.0 | 54.8 | 86.8 | 85.3 | 84.8 | 0.05 | 0.03 | −0.01 | 0.04 | 0.02 | −0.01 | 40.9 | 41.7 | 41.6 |
| CLAP‑Music&Speech | 64.9 | 62.6 | 64.9 | 89.8 | 88.0 | 88.6 | 0.16 | 0.11 | 0.14 | 0.14 | 0.09 | 0.12 | 29.6 | 30.8 | 30.9 |
| Qwen2‑Audio | 58.4 | 58.0 | 59.5 | 88.0 | 86.5 | 86.9 | 0.05 | 0.06 | 0.08 | 0.04 | 0.05 | 0.08 | 36.7 | 37.3 | 37.3 |

Figure 3
Radar Plot Comparison of Top‑Performing Computational Methods. Performance comparison of the eight highest‑ranked methods, averaged across three similarity dimensions, with metrics normalized to [0, 1] scale, where higher values indicate better performance (mean absolute error is inverted).
Table 6
Cross‑Cultural Discrimination Analysis Using Distance‑Based Separation Ratios. Comparison of cultural boundary detection capabilities between humans and all computational methods. Higher values indicate better discrimination between musical traditions. Annotated pairs use only human‑annotated audio pairs (), while all pairs use the complete similarity matrix ( k pairs).
| Method | Annotated Pairs | All Pairs |
|---|---|---|
| Human similarities | ||
| Overall Music | 1.803 | — |
| Cultural | 2.361 | — |
| Recommendation‑level | 2.106 | — |
| Signal‑processing features | ||
| Melody | 1.276 | 1.180 |
| Rhythm | 0.989 | 1.037 |
| Harmony | 1.018 | 1.031 |
| Timbre | 1.025 | 1.018 |
| Foundation models | ||
| MERT‑95 | 1.280 | 1.209 |
| CultureMERT | 1.259 | 1.180 |
| CultureMERT‑TA | 1.262 | 1.169 |
| MERT‑330 | 1.401 | 1.298 |
| CLAP‑Music | 1.415 | 1.217 |
| CLAP‑Music&Speech | 1.366 | 1.318 |
| Qwen2‑Audio | 1.579 | 1.602 |

Figure 4
Linear Regression Weights for Signal Processing Features. Bar charts showing the contribution of signal‑processing features (melody, rhythm, harmony, timbre) in predicting human similarity judgments and foundation model similarities. Positive weights indicate that higher feature similarity contributes to greater predicted similarity, with MAE values in parentheses indicating prediction accuracy.
Table 7
Ensemble Regression Results Combining Signal‑Processing Features and Foundation Models. Performance evaluation of ensemble methods for predicting human similarity judgments. Values are shown as percentages (%) for Triplet Agreement, NDCG, and MAE and as correlation values for Spearman and Kendall metrics. Arrows indicate whether higher () or lower () values represent better performance, with the best performance within each similarity dimension and metric shown in bold.
| Method | Triplet Agr. (%) | NDCG (%) | Spearman (−1, 1) | Kendall (−1, 1) | MAE (%) | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Similarity type | Overall | Cultural | Recomm. | Overall | Cultural | Recomm. | Overall | Cultural | Recomm. | Overall | Cultural | Recomm. | Overall | Cultural | Recomm. |
| Linear regression | 67.0 | 66.7 | 65.1 | 92.5 | 91.4 | 90.9 | 0.19 | 0.15 | 0.18 | 0.18 | 0.14 | 0.17 | 19.7 | 22.2 | 23.0 |
| LightGBM | 67.2 | 63.8 | 64.4 | 92.2 | 90.6 | 90.1 | 0.19 | 0.12 | 0.13 | 0.19 | 0.12 | 0.13 | 19.8 | 22.2 | 23.2 |
