Nalandadata/DrishtiTable
收藏资源简介:
DrishtiTable是一个专门用于表格结构识别(TSR)任务的基准数据集,包含来自印度学术教科书的表格图像及其高质量的HTML结构注释。该数据集由S. Chand Publications出版的教科书中的表格组成,旨在评估模型在特定领域教育内容上的表格结构识别能力。样本发布包含20个表格(10个训练,5个验证,5个测试),完整数据集包含1,421个表格,涵盖6个不同学科的9本教科书。每个样本包括表格图像、HTML注释和元数据。数据集还提供了基准测试结果、统计信息、数据格式和文件结构等详细信息。
DrishtiTable is a curated dataset of table images with high-quality HTML structure annotations from Indian academic textbooks published by S. Chand Publications. It serves as a benchmark for evaluating Table Structure Recognition (TSR) models on domain-specific educational content. The sample release contains 20 representative tables (10 train / 5 val / 5 test), while the full dataset contains 1,421 tables spanning 9 Indian academic textbooks across 6 subjects. Each sample consists of a table image, HTML annotation, and metadata. The dataset also includes benchmark results, statistics, data format, and file structure details.




