遇见数据集

rootsautomation/SciTSR-pd

收藏
Hugging Face2026-05-13 更新2026-06-14 收录
官方服务:

资源简介:

SciTSR-PD是SciTSR数据集的一个公共领域子集,专门用于表格结构识别任务。SciTSR是一个从arXiv LaTeX源文件中提取的大规模科学表格数据集,而SciTSR-PD仅包含源论文采用CC0或等效公共领域许可的表格,这意味着无需署名,也没有商业或衍生使用限制。该子集包含108个表格,来自52篇论文,分为训练集(89个表格)和测试集(19个表格)。每个表格数据包括唯一标识符、论文信息(如标题、作者、许可)、表格图像(PNG格式)、原始PDF、预提取的文本块(带边界框坐标)、单元格结构注释(如ID、内容、行列范围)以及文本块间的关系(仅训练集可用)。数据集主要用于图像到文本和对象检测任务,支持科学表格的文档理解和结构识别研究。注意事项包括注释质量可能因自动生成而有噪声,特别是跨单元格表格,因此建议将注释视为弱监督而非绝对真值。

SciTSR-PD is a public-domain subset of the SciTSR dataset, designed for table structure recognition tasks. SciTSR is a large-scale dataset of scientific tables extracted from arXiv LaTeX source files, and SciTSR-PD includes only tables whose source papers are under CC0 or equivalent public domain licenses, meaning no attribution is required and there are no restrictions on commercial or derivative use. This subset contains 108 tables from 52 papers, split into a training set (89 tables) and a test set (19 tables). Each table entry includes a unique identifier, paper information (e.g., title, authors, license), table image (PNG format), raw PDF, pre-extracted text chunks (with bounding box coordinates), cell structure annotations (e.g., ID, content, row and column spans), and relations between chunks (available only for the training split). The dataset is primarily used for image-to-text and object detection tasks, supporting research in document understanding and structure recognition for scientific tables. Caveats include potential noise in annotations due to automatic generation from LaTeX sources, especially for tables with spanning cells, so annotations should be treated as weak supervision rather than ground truth.

提供机构:
rootsautomation
二维码
社区交流群
二维码
科研交流群
商业服务