nlp-policy-nz-cloud-ocr-pilots
收藏资源简介:
NLP Policy NZ Cloud OCR Pilot Evidence 是一个仅包含元数据的运行报告数据集,旨在支持光学字符识别(OCR)管道操作和基准测试,而非提供完整的OCR语料库。它来源于一个有限的OCR试点项目,具体运行标识为 `track91-zero-cost-hf-20260715`,属于 `cloud-ocr-baseline` 集合,并采用 `metadata_only`(仅元数据)有效载荷策略。报告内容包括运行标识符、来源引用、内容摘要、发布决策和分类计数,但不包含源图像或OCR文本本身。数据集基于Tesseract OCR文档中的示例图像(phototest.tif)生成,用户需遵守源仓库的相关许可和归属条款。该数据集适用于检查有界OCR编排、协调和来源追踪,但不作为已验证的生产基准或法律建议。数据规模为1个发布项,最大配置量为3项,数据文件以JSON格式存储。
NLP Policy NZ Cloud OCR Pilot Evidence is a run-report dataset containing solely metadata, developed to support optical character recognition (OCR) pipeline operations and benchmarking rather than providing a complete OCR corpus. It originates from a limited OCR pilot project, with a specific run identifier `track91-zero-cost-hf-20260715`, belonging to the `cloud-ocr-baseline` collection, and utilizing the `metadata_only` (metadata-only) payload strategy. The report contents include run identifiers, source citations, content summaries, release decisions and classification counts, but exclude the original source images or the OCR text itself. This dataset is generated from the sample images (phototest.tif) included in the Tesseract OCR documentation, and users are required to adhere to the relevant license and attribution terms of the source repository. This dataset is applicable for examining bounded OCR orchestration, coordination and provenance tracking, but shall not be used as a validated production benchmark or legal advice. The dataset comprises 1 release entry, with a maximum configuration count of 3 items, and all data files are stored in JSON format.
数据集概述
- 名称: NLP Policy NZ Cloud OCR Pilot Evidence
- 许可证: 其他(需遵循来源仓库的许可和署名条款)
- 语言: 英语
- 标签: 光学字符识别(OCR)、基准测试、仅元数据、新西兰
- 数据集配置: 默认配置,包含训练集(数据路径:
cloud-ocr-runs/**/*.json)
内容与范围
- 性质: 仅包含元数据的运行报告,来自一次受限的OCR试点项目,用于管道操作和基准测试证据,而非完整的OCR语料库。
- 运行标识:
track91-zero-cost-hf-20260715 - 数据收集:
cloud-ocr-baseline - 负载策略:
metadata_only(仅元数据) - 已发布项目数: 1
- 最大配置容量: 3个项目
- 文档内容: 包含运行标识符、来源引用、内容摘要、发布决策和分类账计数,不包含源图像或OCR文本的通用语料库。
来源与权限
- 试点来源: 来自Tesseract文档的固定示例图像(phototest),地址为
https://raw.githubusercontent.com/tesseract-ocr/tessdoc/a8bf3c9241fe8241f7d13a796f9902c6f0a98087/examples/phototest.tif。 - 使用要求: 用户必须遵循来源仓库的适用许可和署名条款。
- 权利记录: 所有者权利批准已在配套注册证据中记录,本卡片未声明提供者接受或DOI。
预期用途
- 用于检查受限OCR编排、对账和溯源流程。
- 不适用于:已验证的生产级OCR基准测试或法律建议。





