Table Recognition Set (TabRecSet)
收藏资源简介:
Table Recognition Set (TabRecSet) 是一个大规模的数据集,专门为野外环境下的端到端表格识别研究设计。该数据集包含38,177个表格,其中20,415个为英文,17,762个为中文,涵盖了从扫描到相机拍摄的各种场景,如文档、Excel表格、考试试卷和财务发票等。TabRecSet的标注非常完整,包括表格主体空间标注、单元格空间与逻辑标注以及文本内容,用于表格检测、表格结构识别和表格内容识别。此外,数据集使用多边形而非传统的边界框或四边形进行空间标注,更适合野外场景中常见的非规则表格。TabRecSet还包含多种表格形式,如规则和非规则表格(旋转、扭曲等),以及完整的和不完整的边框表格。数据集的应用领域旨在解决端到端表格识别中的挑战,特别是在复杂和多变的野外环境中。
Table Recognition Set (TabRecSet) is a large-scale dataset specifically designed for end-to-end table recognition research in in-the-wild scenarios. It contains 38,177 tables, among which 20,415 are in English and 17,762 are in Chinese, covering a wide range of scenarios from scanned documents to camera-captured materials, such as documents, Excel spreadsheets, exam papers and financial invoices. TabRecSet has comprehensive annotations, including spatial annotations of table bodies, spatial and logical annotations of table cells as well as text contents, which support table detection, table structure recognition and table content recognition tasks. Moreover, the dataset adopts polygon-based spatial annotations instead of traditional bounding boxes or quadrilaterals, which is more suitable for irregular tables commonly encountered in in-the-wild scenarios. Additionally, TabRecSet includes diverse table types, such as regular and irregular tables (e.g., rotated, distorted ones) as well as tables with complete and incomplete borders. This dataset aims to address the challenges in end-to-end table recognition, especially in complex and variable in-the-wild environments.

- 1A large-scale dataset for end-to-end table recognition in the wild华南理工大学电子与信息工程学院 · 2023年



