Nigina-Rinatova/turkic-nlp-corpus
收藏资源简介:
这是一个多语言文本数据集,包含多种突厥语系语言的文本数据,涉及语言有阿塞拜疆语(az)、巴什基尔语(ba)、哈萨克语(kk)、吉尔吉斯语(ky)、土耳其语(tr)、鞑靼语(tt)、维吾尔语(ug)和乌兹别克语(uz)等。数据集配置多样,每个配置具有相同的特征字段:text(文本内容)、lang(语言标签)、script(脚本类型)和source(数据来源)。数据划分主要为训练集(train),部分配置以cc100为后缀,可能基于CC-100语料库构建。数据集规模较大,总示例数量从数千到数百万不等,适用于自然语言处理任务如语言建模、机器翻译等。
This is a multilingual text dataset comprising text data from various Turkic languages. The covered languages include Azerbaijani (az), Bashkir (ba), Kazakh (kk), Kyrgyz (ky), Turkish (tr), Tatar (tt), Uyghur (ug), Uzbek (uz), and others. The dataset features diverse configurations, each sharing the identical set of feature fields: text (text content), lang (language label), script (script type), and source (data source). The dataset is primarily split into the training set (train); some configurations with the "cc100" suffix may be constructed based on the CC-100 corpus. The dataset has a large scale, with the total number of samples ranging from thousands to millions, and it is suitable for natural language processing tasks such as language modeling and machine translation.



