document-review-data
收藏资源简介:
Document Review Data 是一个私有数据集,专为 Office/PDF 文档标题提取审查应用而构建。该数据集包含431个样本,对应431个源文件,其中231个文件为石墨(Shimo)原始XML格式。它主要用于文档标题提取任务的模型训练、验证或人工审查,部署配置文件位于manifests/hf_manifest.jsonl。请注意,该数据集为私有性质,未经许可不得公开,且在使用前需清理源文件。
Document Review Data is a private dataset specifically constructed for Office/PDF document title extraction and review applications. It contains 431 samples corresponding to 431 source files, with 231 files in the original Shimo XML format. The dataset is primarily used for model training, validation, or manual review in document title extraction tasks, with the deployment configuration file located at manifests/hf_manifest.jsonl. Please note that this dataset is private and cannot be publicly disclosed without permission, and source files must be cleaned before use.
数据集概述
- 数据集名称:Document Review Data
- 用途:用于Office/PDF标题提取审核应用的私有数据集
- 样本数量:431个
- 源文件数量:431个
- 其他文件:包含231个Shimo原始XML文件
访问与使用限制
- 该数据集为私有数据集,未经源文件清理不得公开。
- 部署合约文件位于:
manifests/hf_manifest.jsonl。
重要提示
- 数据集详情页地址为:https://huggingface.co/datasets/mannycooper/document-review-data




