遇见数据集

ajdajd/data-snapshot-rev1

收藏
Hugging Face2026-04-28 更新2026-05-03 收录
官方服务:

资源简介:

`data-snapshot`数据集是一个标注语料库,旨在评估和开发从PDF文档中提取数据快照的模型。数据快照被定义为包含来自统计、指标或结构化数据源的定量数据的图表或表格。数据集结构包括注释文件(JSON格式,包含数据快照的对象类别和边界框位置)、原始PDF文件、文档级元数据等。注释文件遵循Data Snapshot Evaluation Format (v1.3)规范。数据集支持英语、法语和西班牙语,主要用于对象检测和图像分割任务。数据来源包括UNHCR等机构。

The `data-snapshot` dataset is an annotated corpus designed for the evaluation and development of models for extracting *data snapshots* from PDF documents. A **data snapshot** is defined as a figure or table that contains quantitative data derived from statistics, indicators, or structured data sources. The dataset includes annotation files (in JSON format, indicating the data snapshots object class and bounding box locations), raw PDFs, document-level metadata, etc. The annotation files follow the Data Snapshot Evaluation Format (v1.3) schema. The dataset supports English, French, and Spanish, and is primarily used for object detection and image segmentation tasks. Data sources include UNHCR and others.

提供机构:
ajdajd
二维码
社区交流群
二维码
科研交流群
商业服务