遇见数据集

ParlMix-UA-RU - Ukrainian Parliamentary Code-Mixing Dataset

收藏
Zenodo2026-05-17 更新2026-05-26 收录
官方服务:

资源简介:

ParlMix-UA-RU: Ukrainian Parliamentary Code-Mixing Dataset Overview ParlMix-UA-RU is a specialized, manually annotated linguistic resource designed to facilitate research on code-mixing (CM) and language identification (LID) between Ukrainian and Russian. The dataset is derived from official transcripts of Ukrainian parliamentary sessions (Verkhovna Rada), capturing language dynamics in a formal, high-register political discourse. Dataset Composition & Selection The corpus was developed through a multi-stage process to ensure its utility for complex Natural Language Processing (NLP) tasks: Code-Mixed Sentences: Initially, sentences were extracted based on a threshold of more than two out-of-vocabulary (OOV) tokens relative to monolingual Ukrainian dictionary. This selection typically indicates intra-sentential code-switching or mixed language use. Monolingual Balancing: To improve the robustness of Language Identification models and prevent classification bias, the dataset was subsequently enriched with monolingual Russian sentences from the parlamentary scripts. This ensures a balanced distribution of tokens across the target languages. Annotation Methodology The dataset follows a rigorous "human-in-the-loop" approach to establish a Gold Standard for linguistic research: Manual Tagging: Every token has been manually labeled with its specific language tag by expert linguists. Zero AI Involvement: No automated labeling or AI systems were used for the final annotations, ensuring maximum precision and the preservation of subtle linguistic nuances. Volume: The final dataset contains approximately 150,000 tokens with token-level manual annotations. Research Applications ParlMix-UA-RU is a valuable resource for: Training and benchmarking Language Identification (LID) systems, especially for closely related languages. Computational analysis of intra-sentential code-mixing patterns. Sociolinguistic studies of language use in official governmental settings. Technical Details Language(s): Ukrainian (UA), Russian (RU) Source: Verkhovna Rada transcripts Format: json Annotation level: Token-level language identification Citation & Attribution If you utilize this resource in your academic work, please use the following citation: Kanishcheva, O., Shvedova, M., Dyka, L., & Husenko, K. (2026). ParlMix-UA-RU: Ukrainian Parliamentary Code-Mixing Dataset (Version 1.0.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.14724542

提供机构:
Zenodo
创建时间:
2025-10-20
二维码
社区交流群
二维码
科研交流群
商业服务