StreamLingDiv
收藏资源简介:
Music, similarly to language, is a prominent carrier of culture. Most music also has lyrics, i.e., words that convey meaning in a certain language. With the rise of streaming services such as Spotify, music consumption has, to a large extent, become digitized, and the music we listen to is shaped by algorithms. Therefore, the language of the music we stream constitutes one dimension of digital language use and digital linguistic diversity. To quantify this dimension and situate it in space and time, we have developed StreamLingDiv. The dataset aggregates the Spotify top 200 charts of streamed songs per country and year from MGD+ (Seufitelli et al, 2023) and combines it with lyrics data from Song Genius available at Kaggle. We run two language identification models on the preprocessed lyrics (FastText and GlotLID) to get language tags and derive country-year-language distribution of streams. For each country and year, we calculate measures of diversity according to the Leinster-Cobbold Framework as done in Essfors (2025) using the ASJP (Wichmann et al., 2025). The dataset covers the years and countries of MGD+, i.e., 2017 - 2022 and 67 countries. 2022 only has data until March, i.e., it risks missing some seasonal variation in music trends, e.g., Christmas music, and should be treated with caution. Also, only 55 countries have data for all years. Because of mismatches in artist and title between the datasets we combine, 65,912 of 110,753 ~60% songs have lyrics whose language can be identified. All in all, we automatically identify 102 languages covering 77 percent of all streams. Cases where the language could not be identified are given the label unknown. We only use labels where the language identification models signal high certainty. Based on the country-year-language distributions, we derive formal measures of linguistic diversity according to the Leinster-Cobbold framework as described in Essfors (2025). Note, we have not made any manual adjustment to the language tags or attempts at outlier detection in this version of the dataset. Based on ad hoc inspection of the language tags, we deem them sufficiently reliable for course-grained analysis. The code to replicate the construction of the dataset is published here: https://github.com/Eszettfors/StreamLingDiv_code. Please note that we do not publish any lyrics data so as not to infringe on any copyrights. We publish four .csv files:- country_lang_streams.csv: This file contains the distribution of streams across languages identified by ISO6393-codes across countries identified by two-letter codes for each year in the data. - country_data.csv: This file contains country metadata accessed from Natural Earth to aid geographical analysis. - language_data.csv: This file contains language metadata from Glottolog, such as glottocodes, language names and language family, to aid linguistic analysis. - diversity_measures.csv: This file contains precalculated measures of diversity based on the language distributions in country_lang_streams.csv excluding "unknown". To quantify uncertainty, both the total number of streams and the percentage of streams for which the language could be identified are supplied as separate columns. Six measures of diveristy is provided: the classical Richness, Exponent Shannon, and Inverse Simpson, together with their lexical similarity-aware implementations. See Essfors (2025) for details. ------------ This research was funded by WWTF (grant number ICT23-012). It is a part of the DIGILINGDIV-project. ------------ Seufitelli, Danilo B.; Oliveira, Gabriel P.; Silva, Mariana O.; Moro, Mirella M. MGD+: An Enhanced Music Genre Dataset with Success-based Networks. In : DATASET SHOWCASE WORKSHOP (DSW), 5th, 2023, Belo Horizonte/MG. Proceedings [...]. Porto Alegre: Brazilian Computer Society, 2023. p. 36-47. DOI: https://doi.org/10.5753/dsw.2023.233826 . Wichmann, Søren, Eric W. Holman, Cecil H. Brown, Matthew S. Dryer, and Qibin Ran (eds.). 2025. The ASJP Database (version 21). Essfors, Hannes. 2025. “Global Linguistic Diversity - Adapting the Leinster-Cobbold Framework from Ecology for Humanities Research” edited by T. Arnold, M. Fantoli, and and R. Ros. Anthology of Computers and the Humanities 3:653–69. doi:10.63744/srhQaCwGo5mj.



