WSHAPER/xdailydialog-ru
收藏资源简介:
XDailyDialog Russian (xdailydialog-ru) 是XDailyDialog英文分割的俄语翻译数据集,通过GLM-5.1模型进行机器翻译,并经过人工质量审查。该数据集用于文本分类任务,特别是对话行为分类,包含俄语对话数据。数据格式每行包括一个对话(话语之间用__eou__分隔)、主题(整数)、每个话语的对话行为(1=承诺,2=指令,3=信息,4=问题)和情感(0=中性,1=愤怒,2=厌恶,3=恐惧,4=快乐,5=悲伤,6=惊讶)。数据集分为训练集(11,110个对话,83,035个话语)、开发集(1,037个对话,8,025个话语)和测试集(1,047个对话,7,716个话语),总规模在10万到100万之间。翻译过程采用句子级翻译,基于哈希缓存和可恢复管道,并进行长度比验证、语言检测和标点保留等质量检查,覆盖率达到100%(98,776个话语全部翻译)。数据集遵循Apache-2.0许可证。
Russian translation of the XDailyDialog English split, produced with GLM-5.1 via machine translation with human-quality review. This dataset is designed for text classification tasks, specifically dialogue act classification, and contains Russian dialogue data. Each line follows the format: a dialogue with utterances separated by __eou__, a topic (integer), dialogue acts per utterance (1=commissive, 2=directive, 3=inform, 4=question), and emotions per utterance (0=neutral, 1=anger, 2=disgust, 3=fear, 4=happiness, 5=sadness, 6=surprise). It includes splits for training (11,110 dialogues, 83,035 utterances), development (1,037 dialogues, 8,025 utterances), and testing (1,047 dialogues, 7,716 utterances), with a total size between 100K and 1M. Translation was performed sentence-by-sentence with hash-based caching and resumable pipeline, validated by length ratios, language detection, and punctuation preservation, achieving 100% coverage (98,776 utterances translated). The dataset is licensed under Apache-2.0.



