bevaya/SciTSR-cc-by-nc-sa
收藏资源简介:
SciTSR-CC-BY-NC-SA是一个经过许可证过滤的数据集子集,源自SciTSR大规模表格结构识别数据集。该数据集包含从arXiv LaTeX源文件中提取的科学表格,并筛选出与CC-BY-NC-SA 4.0开源模型发布兼容的许可证(包括公共领域、CC-BY、CC-BY-NC和CC-BY-NC-SA许可证的论文)。数据集本身以CC-BY-NC-SA 4.0许可证发布。它包含889个表格(训练集697个,测试集192个),来自470篇论文,涵盖表格图像、PDF、文本块坐标、单元格结构注释和关系标签,主要用于表格结构识别和文档理解任务,适用于非商业研究。
A license-filtered subset of SciTSR, a large-scale table structure recognition dataset of scientific tables extracted from arXiv LaTeX source files. This subset contains tables whose source papers are compatible with a CC-BY-NC-SA 4.0 open-weight model release — covering public domain, CC-BY, CC-BY-NC, and CC-BY-NC-SA licensed papers. The dataset itself is released under CC-BY-NC-SA 4.0. It includes 889 tables (697 train, 192 test) from 470 papers, with features such as table images, PDFs, text chunk coordinates, cell structure annotations, and relation labels, designed for table structure recognition and document understanding tasks in non-commercial research.



