遇见数据集

rootsautomation/SciTSR-cc-by-nc-sa

收藏
Hugging Face2026-05-13 更新2026-06-14 收录
官方服务:

资源简介:

SciTSR-CC-BY-NC-SA 是 SciTSR 数据集的一个经过许可证过滤的子集,专门用于表格结构识别任务。该数据集包含从 arXiv LaTeX 源文件中提取的科学表格,经过筛选仅保留与 CC-BY-NC-SA 4.0 开源模型发布兼容的表格(包括公共领域、CC-BY、CC-BY-NC 和 CC-BY-NC-SA 许可证的论文)。数据集本身以 CC-BY-NC-SA 4.0 许可证发布。总共有 889 个表格,来自 470 篇论文,分为训练集(697 个表格)和测试集(192 个表格)。每个表格行包含图像渲染、PDF 文件、文本块、单元格结构注释和关系标签,适用于图像到文本、目标检测等任务,但主要用于非商业用途。数据集旨在支持文档理解和科学表格分析,同时注意标注质量可能存在噪声,且需遵守原始作者的署名要求。

A license-filtered subset of SciTSR, a large-scale table structure recognition dataset of scientific tables extracted from arXiv LaTeX source files. This subset contains tables whose source papers are compatible with a CC-BY-NC-SA 4.0 open-weight model release — covering public domain, CC-BY, CC-BY-NC, and CC-BY-NC-SA licensed papers. The dataset itself is released under CC-BY-NC-SA 4.0. It includes 889 tables from 470 papers, split into train (697 tables) and test (192 tables). Each row provides an image render, PDF, text chunks, cell structure annotations, and relation labels, designed for tasks like image-to-text and object detection, primarily for non-commercial use in document understanding and scientific table analysis.

提供机构:
rootsautomation
二维码
社区交流群
二维码
科研交流群
商业服务