adorkin/olmocr_science_pdfs-history_and_geography
收藏资源简介:
该数据集是一个文本数据集,名为dolma3_pool,具体来自olmocr_science_pdfs-history_and_geography子集。它包含1,674,875个训练样本,每个样本具有三个特征:id(字符串类型,表示唯一标识符)、text(字符串类型,表示文本内容)和n_tokens(整数类型,表示文本中的令牌数量)。数据集总大小为107,541,463,905字节,下载大小为58,932,364,461字节。数据文件以train-*格式存储在指定路径中,适用于自然语言处理任务,如文本分析和模型训练。
This dataset is a text dataset named dolma3_pool, specifically from the olmocr_science_pdfs-history_and_geography subset. It contains 1,674,875 training examples, each with three features: id (string type, representing a unique identifier), text (string type, representing the text content), and n_tokens (integer type, indicating the number of tokens in the text). The total dataset size is 107,541,463,905 bytes, with a download size of 58,932,364,461 bytes. The data files are stored in the specified path in the format train-*, suitable for natural language processing tasks such as text analysis and model training.



