Investigating genre similarity boundaries in Brazilian music using deep learning models
DOI:
https://doi.org/10.5216/mh.v26.85233Palabras clave:
music information retrieval, music genre recognition, Brazilian music, vision transformers, embeddingsResumen
Understanding how musical genres are organized in learned representation spaces remains a central challenge in music information retrieval. This study investigates the latent structure induced by a Vision Transformer (ViT) model fine-tuned on Brazilian regional music. We introduce the Brazilian Regional Music Dataset (BYRM), a curated collection of 1,082 tracks distributed across ten culturally diverse genres. Our analysis focuses on the best-performing experimental configuration identified in prior experiments, in which 10-second audio segments are extracted from the 90–120 second portion of each track. Mel-spectrogram representations are used as model input, and time-local embeddings are derived from the trained ViT. To examine the organization of the learned feature space, we apply dimensionality reduction techniques, including Principal Component Analysis (PCA), t-distributed Stochastic Neighbor Embedding (t-SNE), and Uniform Manifold Approximation and Projection (UMAP). Cosine similarity is computed to quantify inter-genre proximity. The model achieves 81.94% classification accuracy and an F1-score of 81.84%, demonstrating strong discriminative capability. Beyond classification performance, the representation analysis reveals coherent clustering patterns, with stylistically related genres exhibiting higher proximity while structurally distinct genres remain well separated. These findings suggest that transformer-based audio embeddings effectively encode culturally grounded genre relationships within a musically diverse context.







