CCNet
收藏资源简介:
CCNet是由Facebook AI创建的一个大规模单语数据集,旨在从Common Crawl中提取高质量文本数据。该数据集包含15亿文档,覆盖174种语言,其中英语文档达到7亿,总Tokens数为5320亿。创建过程中,采用了文档去重和语言识别技术,并通过与高质量数据源如Wikipedia的相似度筛选文档。CCNet主要用于训练文本表示模型,特别是在低资源语言上,以提高自然语言处理任务的性能。
CCNet is a large-scale monolingual dataset created by Facebook AI, designed to extract high-quality textual data from Common Crawl. This dataset encompasses 1.5 billion documents spanning 174 languages, among which English documents reach 700 million, with a total token count of 532 billion. During its curation, document deduplication and language identification techniques were adopted, and documents were screened based on their similarity to high-quality data sources such as Wikipedia. CCNet is primarily utilized for training text representation models, especially for low-resource languages, to improve the performance of natural language processing tasks.



