ParlMix-UA-RU - Ukrainian Parliamentary Code-Mixing Dataset
收藏官方服务:
资源简介:
The UkrRusParlMix dataset contains sentences extracted from Ukrainian parliamentary transcripts to support research on code-mixing between Ukrainian and Russian. Only sentences with more than two out-of-vocabulary words were kept, typically indicating mixed language use. Each token was manually labeled with its language, resulting in a balanced dataset of about 150,000 tokens suitable for language identification and code-mixing analysis.
创建时间:
2025-01-23



