遇见数据集

PRESEEA-Málaga I: A Standardized Sociolinguistic Database Linking 99,000+ Phonetic and Lexical Annotations to Speaker Profiles

收藏
Zenodo2025-12-13 更新2026-05-26 收录
官方服务:

资源简介:

This paper presents a sociolinguistic data set that aggregates sociological, phonetic and lexical data from speakers of Andalusian Spanish that allows to asses the capability of Machine Learning models to predict their sociological profile by their linguistic characteristics. Earlier analyses of this corpus collected in the mid-1990s (PRESEEA-Málaga) were hampered by heterogeneous data formats and inconsistent codification. Here, we provide a resource built as a tidy-data SQLite database that presents each variable into separate, well‑documented tables, standardized categorical levels, and encoded missing values explicitly. The resulting schema integrates more than sixty sociological attributes (e.g., sex, age, education, socioeconomic situation, media habits, parental data, network ties) with detailed phonetic annotations for six pronunciation variables and variation analysis of six lexical items. In total, the linguistics analysis provides nearly 100,000 human-expert labeled datapoints. In addition to the SQLite database, we provide .xlsx, .csv and RData files and scripts in both R and Python for general purpose post-processing. There is a manuscript for the paper which provides full information about the dataset, however, it has yet to be published. We will update this description once it happens, however, for now, if your require additional information, feel free to reach out to us.

提供机构:
Zenodo
创建时间:
2025-12-13
二维码
社区交流群
二维码
科研交流群
商业服务