A Chinese-Uighur comparable corpus
收藏Mendeley Data2024-01-31 更新2024-06-28 收录
下载链接:
https://www.scidb.cn/en/detail?dataSetId=752811851674288128
下载链接
链接失效反馈官方服务:
资源简介:
The dataset is composed of comparable corpus of Chinese and Uigur, obtained from the Internet. Chinese and Uigur language pairs are textually corresponding. The dataset is mainly from news, including news headlines, time and text. The dataset contains two data files: ch_corpus.zip and uy_corpus.zip. Each package contains four documents, namely document_1, document_2, document_3 and document_4. Each document contains two folders: uy and ch, where uy represents Uyghur, ch represents Chinese, and each folder contains multiple text documents. Uighur and Chinese language pairs are organized correspondingly according to their names.
创建时间:
2024-01-31



