A Chinese-Uighur comparable corpus

Mendeley Data2024-01-31 更新2024-06-28 收录

下载链接：

https://www.scidb.cn/en/detail?dataSetId=752811851674288128

下载链接

链接失效反馈

官方服务：

资源简介：

The dataset is composed of comparable corpus of Chinese and Uigur, obtained from the Internet. Chinese and Uigur language pairs are textually corresponding. The dataset is mainly from news, including news headlines, time and text. The dataset contains two data files: ch_corpus.zip and uy_corpus.zip. Each package contains four documents, namely document_1, document_2, document_3 and document_4. Each document contains two folders: uy and ch, where uy represents Uyghur, ch represents Chinese, and each folder contains multiple text documents. Uighur and Chinese language pairs are organized correspondingly according to their names.

创建时间：

2024-01-31

5,000+

优质数据集

54 个

任务类型

进入经典数据集