CCpdf
收藏资源简介:
CCpdf数据集是由雪花科技和亚当密茨凯维奇大学合作创建的,旨在从互联网上的PDF文件中构建一个大规模、多样化的多语言文档语料库。该数据集包含1450万页PDF文件,覆盖11种不同语言,主要来源于2010年至2022年间的文档。创建过程中,研究团队分析了多种处理技术,以平衡数据质量、处理时间和成本。CCpdf数据集特别适用于2D语言模型的预训练,有助于提升模型在多语言和多领域文档理解方面的性能。
The CCpdf dataset was jointly developed by Snowflake Technology and Adam Mickiewicz University, with the objective of constructing a large-scale, diverse multilingual document corpus from PDF documents sourced from the public Internet. This dataset comprises 14.5 million pages of PDF documents, spanning 11 distinct languages, and is primarily derived from documents published between 2010 and 2022. During its development, the research team evaluated multiple processing techniques to strike a balance between data quality, processing latency, and associated costs. The CCpdf dataset is particularly suitable for pre-training 2D language models, and helps improve the model's performance in multilingual and multi-domain document understanding.

- 1CCpdf: Building a High Quality Corpus for Visually Rich Documents from Web Crawl Data雪花科技 · 2023年



