遇见数据集

A Chinese-Uighur comparable corpus

收藏
科学数据银行2020-09-08 更新2026-04-23 收录
官方服务:

资源简介:

The dataset is composed of comparable corpus of Chinese and Uigur, obtained from the Internet. Chinese and Uigur language pairs are textually corresponding. The dataset is mainly from news, including news headlines, time and text. The dataset contains two data files: ch_corpus.zip and uy_corpus.zip. Each package contains four documents, namely document_1, document_2, document_3 and document_4. Each document contains two folders: uy and ch, where uy represents Uyghur, ch represents Chinese, and each folder contains multiple text documents. Uighur and Chinese language pairs are organized correspondingly according to their names.

创建时间:
2020-09-08
二维码
社区交流群
二维码
科研交流群
商业服务