遇见数据集

adorkin/olmocr_science_pdfs-literature

收藏
Hugging Face2026-05-01 更新2026-05-31 收录
官方服务:

资源简介:

该数据集名为dolma3_pool,具体为olmocr_science_pdfs-literature子集,由AllenAI提供。它包含训练分片,共有1,846,289个示例,每个示例具有id、text和n_tokens三个特征,其中id为字符串类型,text为字符串类型,n_tokens为整数类型。数据集总大小为128,398,816,781字节,下载大小为72,895,677,988字节。数据文件路径为data/train-*,主要用于自然语言处理任务,可能涉及科学PDF文献的OCR文本处理。

The dataset is named dolma3_pool, specifically the olmocr_science_pdfs-literature subset, provided by AllenAI. It includes a training split with 1,846,289 examples, each featuring id, text, and n_tokens, where id is of string type, text is of string type, and n_tokens is of integer type. The total dataset size is 128,398,816,781 bytes, with a download size of 72,895,677,988 bytes. The data files are located at data/train-*, and it is primarily used for natural language processing tasks, potentially involving OCR text processing of scientific PDF literature.

提供机构:
adorkin
二维码
社区交流群
二维码
科研交流群
商业服务