CLUECorpus2020
收藏资源简介:
CLUECorpus2020是由CLUE组织创建的大型中文语料库,旨在支持语言模型的预训练和语言生成。该数据集包含100GB的原始文本,总计350亿个中文字符,来源于Common Crawl。数据集被分为训练、开发和测试集,每个文件都遵循预训练格式。创建过程中,通过详细的过滤和提取规则,确保数据质量。CLUECorpus2020广泛应用于中文自然语言处理任务,如语言理解和生成,旨在提升模型在中文环境下的性能。
CLUECorpus2020 is a large-scale Chinese corpus created by the CLUE organization, designed to support the pre-training of language models and language generation. The dataset contains 100GB of original text, totaling 35 billion Chinese characters, sourced from Common Crawl. It is divided into training, development, and test sets, with each file adhering to the pre-training format. During the creation process, detailed filtering and extraction rules were applied to ensure data quality. CLUECorpus2020 is widely used in Chinese natural language processing tasks such as language understanding and generation, aiming to enhance model performance in the Chinese environment.




