olmOCR-mix-1025-Photoreal
收藏资源简介:
本数据集是allenai/olmOCR-mix-1025数据集的派生作品,专注于对文档页面进行真实感增强。通过Synthetic Engine,将干净的PDF渲染转换为具有自然光照、纸张纹理、阴影和拍摄伪影的真实扫描或照片图像,用于模拟现实世界中的文档捕获条件。图像文件采用JPEG格式,文件名与原始数据集一致,用户需自行从原始数据集下载对应的纯文本注释(无边界框或行级多边形)。当前仅包含100页咖啡渍效果的增强图像,未来将持续添加更多效果。数据集规模小于1000个样本,适用于图像到文本(OCR)和文本生成任务,旨在提升模型在真实场景下的鲁棒性。
This dataset is a derived work of the allenai/olmOCR-mix-1025 dataset, focusing on realistic enhancement of document pages. Through the Synthetic Engine, clean PDF renders are transformed into real scans or photo images with natural lighting, paper texture, shadows, and capture artifacts, simulating real-world document capture conditions. Images are in JPEG format with filenames consistent with the original dataset; users need to download the corresponding plain text annotations (without bounding boxes or line-level polygons) from the original dataset. Currently, it contains only 100 enhanced images with coffee stain effects, with more effects to be added in the future. The dataset size is under 1000 samples, suitable for image-to-text (OCR) and text generation tasks, aiming to improve model robustness in real-world scenarios.
数据集概述:AlroWilde/olmOCR-mix-1025-Photoreal
基本信息
- 许可证:ODC-BY(开放数据共享署名许可)
- 任务类别:文本生成(text-generation)、图像到文本(image-to-text)
- 数据规模:少于1K条样本(n<1K)
数据集简介
该数据集是 allenai/olmOCR-mix-1025 的衍生作品,对其文档页面进行了写实增强处理。原始干净的PDF渲染图被转换为具有自然光照、纸张纹理、阴影和拍摄伪影的真实感扫描/拍摄图像,该增强处理由 Synthetic Engine 工具完成。
文件与注释说明
- 图像文件名与原始数据集完全一致,仅扩展名改为
.jpeg。 - 不提供额外的parquet注释文件。原始注释为纯文本格式(不含边界框或行级多边形),需从原始数据集中按文件名匹配下载。
当前增强效果
- 咖啡渍效果(Coffee stains):100页(添加日期:2026-07-28)
- 新效果将持续不断添加,后续会有更多增强类型上线。
设计目的
该数据集旨在为真实世界条件下训练更鲁棒的OCR(光学字符识别)和文档理解模型而设计。
使用与反馈
- 数据集作者欢迎用户反馈该数据集对模型训练带来的正面或负面改进效果,可通过
hi@support.alrowilde.com邮箱联系。 - 如需新的增强效果或有反馈,可在数据集页面留言或发送邮件。
衍生声明
该数据集是 allenai/olmOCR-mix-1025(ODC-BY许可证)的衍生作品,使用时需按要求标注原始数据集。





