DCAD-2000
收藏资源简介:
DCAD-2000是一个大规模、高质量的多语言数据集,由清华大学等机构构建,包含2282种语言,数据量达到8.63亿份文档,46.72TB的存储大小,涵盖了155种高和中资源语言以及159种书写脚本。该数据集通过将数据清洗任务重新定义为异常检测问题,显著提高了数据质量,能够识别并移除噪声或异常内容。DCAD-2000适用于多种下游自然语言处理任务,特别是在提高低资源语言的多语言模型性能方面表现出色。
DCAD-2000 is a large-scale, high-quality multilingual dataset constructed by Tsinghua University and other institutions. It covers 2282 languages, with a total of 863 million documents and a storage size of 46.72 TB, and includes 155 high- and medium-resource languages as well as 159 writing scripts. By redefining the data cleaning task as an anomaly detection problem, this dataset significantly improves data quality by identifying and removing noisy or anomalous content. DCAD-2000 is applicable to a wide range of downstream natural language processing tasks, and particularly excels in enhancing the performance of multilingual models for low-resource languages.

- 1DCAD-2000: A Multilingual Dataset across 2000+ Languages with Data Cleaning as Anomaly Detection清华大学 · 2025年



