遇见数据集

smatiush/pdfa-eng-wds

收藏
Hugging Face2026-05-07 更新2026-05-31 收录
官方服务:

资源简介:

PDFA数据集是一个从SafeDocs语料库(即CC-MAIN-2021-31-PDF-UNTRUNCATED)过滤而来的文档数据集。原始语料库的目的是进行全面的PDF文档分析,而该子集的目的是为视觉语言模型提供机器学习就绪的数据。该数据集以webdataset .tar格式提供,包含PDF文件和包含OCR注释及元数据的JSON文件配对。数据集经过过滤,移除了大于100MB或渲染时间超过500ms的文件,并使用XLM-Roberta模型将数据限制为英语子集。每个文档的元数据包括单词、边界框、行信息以及嵌入图像的边界框,数据格式针对大规模多模态机器学习(特别是图像到文本任务)进行了优化。数据集包含2,159,432个样本,约1800个分片,总计约18M页和9.7B个令牌。

The PDFA dataset is a document dataset filtered from the SafeDocs corpus (i.e., CC-MAIN-2021-31-PDF-UNTRUNCATED). The original corpus was intended for comprehensive PDF document analysis, while this subset is designed to provide machine learning-ready data for vision-language models. This dataset is provided in the webdataset .tar format, containing paired PDF files and JSON files with OCR annotations and metadata. The dataset has been filtered to remove files larger than 100 MB or with rendering time exceeding 500 ms, and restricted to the English subset using the XLM-Roberta model. The metadata for each document includes words, bounding boxes, line information, and bounding boxes of embedded images, with the data format optimized for large-scale multimodal machine learning, particularly image-to-text tasks. The dataset contains 2,159,432 samples across approximately 1,800 shards, totaling around 18 million pages and 9.7 billion tokens.

提供机构:
smatiush
二维码
社区交流群
二维码
科研交流群
商业服务