ajdajd/data-snapshot-rev0
收藏资源简介:
`data-snapshot`数据集是一个标注语料库,旨在评估和开发从PDF文档中提取*数据快照*的模型。**数据快照**被定义为包含来自统计、指标或结构化数据源的定量数据的图表或表格。数据集的结构包括注释文件(指示数据快照的对象类别和边界框位置)、原始PDF文件、文档级元数据和一个注释文件的模式。注释文件遵循Data Snapshot Evaluation Format (v1.3),并通过人类标注使用Label Studio完成。数据来源包括UNHCR、PRWP(WIP)和Refugee(WIP)。
The `data-snapshot` dataset is an annotated corpus designed for the evaluation and development of models for extracting *data snapshots* from PDF documents. A **data snapshot** is defined as a figure or table that contains quantitative data derived from statistics, indicators, or structured data sources. The dataset structure includes annotation files (indicating the object class and bounding box locations of data snapshots), raw PDF files, document-level metadata, and a schema for the annotation files. The annotation files follow the Data Snapshot Evaluation Format (v1.3) and were produced through human labeling using Label Studio. Sources include UNHCR, PRWP (WIP), and Refugee (WIP).



