synth_shamela_ocr_arabic_books
收藏资源简介:
该数据集是一个包含100本独立书籍数字化页面的集合,以多配置形式组织(例如每本书对应一个配置如book_1)。每个样本代表书籍的一个页面,包含四个核心字段:image存储页面图像数据,page_header记录页眉文本,body记录正文文本,footnotes记录脚注文本。所有数据仅划分为训练集,采用CC BY-NC-ND 4.0许可证。这是一个多模态(图像与文本)数据集,适用于文档图像分析、光学字符识别(OCR)、版面结构识别、历史文献数字化以及相关自然语言处理与计算机视觉任务的研究与模型训练。
This dataset is a collection of digitized pages from 100 standalone books, organized in multiple configurations (`config`), where each book corresponds to one configuration (e.g., `book_1`, `book_9`). Each sample represents a single page of a book and contains four core fields: the `image` field stores the image data of the page; the `page_header` field records the page's header text as a string; the `body` field records the page's main body text as a string; the `footnotes` field records the page's footnote text as a string. All data is exclusively split into the training set (`train` split). This dataset is licensed under CC BY-NC-ND 4.0. As a multimodal (image and text) dataset, it is applicable to research and model training for tasks including document image analysis, optical character recognition (OCR), layout structure recognition, historical document digitization, and related natural language processing and computer vision tasks.
数据集概要
- 数据集名称:synth_shamela_ocr_arabic_books
- 许可证:CC BY-NC-ND 4.0
- 数据集地址:https://huggingface.co/datasets/freococo/synth_shamela_ocr_arabic_books
数据集配置
该数据集包含多个子配置(config_name),每个配置对应一本书籍。具体配置名称如下:
- book_1
- book_9
- book_11
- book_12
- book_14
- book_20
- book_25
- book_28
- book_33
- book_35
- book_36
- book_37
- book_39
- book_40
- book_42
- book_46
- book_51
- book_56
- book_59
- book_63
- book_64
- book_65
- book_69
- book_70
- book_71
- book_73
- book_75
- book_76
- book_79
- book_81
- book_83
- book_90
- book_91
- book_92
- book_93
- book_95
- book_98
- book_110
- book_111
- book_112
- book_115
- book_117
- book_118
- book_136
- book_140
- book_144
- book_145
- book_151
- book_152
- book_154
- book_160
- book_161
- book_162
- book_169
- book_170
- book_172
- book_173
- book_174
- book_175
- book_177
- book_178
- book_183
- book_195
- book_196
- book_211
- book_212
- book_214
- book_218
- book_219
- book_225
- book_232
- book_233
- book_234
- book_248
- book_250
- book_252
- book_254
- book_260
- book_263
- book_266
- book_272
数据集特征
每个配置均包含以下四个特征字段:
- image(图像类型):书籍页面的图像数据。
- page_header(字符串类型):页眉文本。
- body(字符串类型):正文文本。
- footnotes(字符串类型):脚注文本。
数据拆分
每个配置仅包含一个训练集(splits: train),无其他拆分。





