遇见数据集

TabbyXL: Experiment Data

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

The data are designed to evaluate TabbyXL, a system for rule-based transformation spreadsheet data from arbitrary to relational tables that is freely available at GitHub (https://github.com/cellsrg/tabbyxl). Our data are based on the existing dataset of tables Troy_200 [1]. It contains 200 arbitrary tables as CSV files collected from 10 different government statistical websites. They were collected for the experiment on data extraction from tables that is presented in the paper [2]. We use its earlier version that stores the original tables with style features (fonts, alignment, and indentation) as Excel spreadsheets available at http://tango.byu.edu/data. We have put all of these tables with style features into the single spreadsheet file (data/TangoDataset.xlsx). Each of 200 tables is located in a separate sheet. The pair of tags $START and $END points out to its location inside the sheet. We initially used this file in our previous experiment described in the paper [3]. We have transformed automatically all tables of the single spreadsheet into the relational form, using TabbyXL and the ruleset (data/rules.dslr). The folder data/results contains the obtained results. The folder data/gt contains the ground-truth data for automated performance evaluation of TabbyXL in the role and structural stages of the table analysis. Each table of our data/results and data/gt dataset is accompanied with two recordsets: ENTRIES and LABELS. The first of them specifies entries. Each record presents an entry as a triple <value, provenance, set of associated labels>. In LABELS recordset each record presents a label as a triple <value, provenance, parent reference>. We also have stored the log files: results.log with the results of running and eval.log with the results of performance evaluation of TabbyXL. REFERENCES [1] Nagy G. TANGO-DocLab web tables from international statistical sites, (Troy_200), 1, ID: Troy_200_1. URL: http://tc11.cvc.uab.es/datasets/Troy_200_1. [2] Embley D., Krishnamoorthy M., Nagy G., & Seth S. (2016). Converting heterogeneous statistical tables on the web to searchable databases. Int. J. on Document Analysis and Recognition, 19(2), 119-138. URL: https://link.springer.com/article/10.1007/s10032-016-0259-1. [3] Shigarov A., Paramonov V., Belykh P., & Bondarev A. (2016) Rule-Based Canonicalization of Arbitrary Tables in Spreadsheets. Proc. 22nd Int. Conf. on Information and Software Technologies, pp. 78-91. URL: http://link.springer.com/chapter/10.1007/978-3-319-46254-7_7.

本数据集用于评估TabbyXL——一款可将任意格式电子表格数据转换为关系型表格的基于规则的系统,该系统可在GitHub(https://github.com/cellsrg/tabbyxl)免费获取。本数据集基于已公开的Troy_200表格数据集[1],包含从10个不同政府统计网站采集的200个任意格式表格,均以CSV文件形式存储。这些表格原本是为文献[2]中提及的表格数据提取实验所采集。本研究使用其早期版本,该版本将带有字体、对齐方式及缩进等格式特征的原始表格存储为Excel电子表格,相关资源可通过http://tango.byu.edu/data获取。 我们已将所有带格式特征的表格整合至单个电子表格文件(data/TangoDataset.xlsx)中,200个表格分别位于独立的工作表内。标签对$START与$END用于标记表格在工作表中的具体位置。本数据集最初应用于文献[3]中描述的前期实验。 我们借助TabbyXL与规则集(data/rules.dslr),将单个电子表格内的所有表格自动转换为关系型格式。data/results文件夹存储转换得到的结果,data/gt文件夹则存储用于TabbyXL表格分析角色与结构阶段自动化性能评估的基准真值(ground-truth)数据。 data/results与data/gt数据集中的每个表格均配套两个记录集:ENTRIES与LABELS。其中ENTRIES记录集用于定义条目,每条记录以三元组<值,来源,关联标签集>的形式呈现一个条目;LABELS记录集中的每条记录则以三元组<值,来源,父级引用>的形式呈现一个标签。 本数据集还存储了两类日志文件:results.log记录系统运行结果,eval.log记录TabbyXL的性能评估结果。 参考文献 [1] Nagy G. 国际统计网站网页表格数据集TANGO-DocLab(Troy_200),1,编号:Troy_200_1。网址:http://tc11.cvc.uab.es/datasets/Troy_200_1。 [2] Embley D., Krishnamoorthy M., Nagy G., Seth S. (2016). 将Web上异构统计表格转换为可搜索数据库. 国际文档分析与识别期刊(Int. J. on Document Analysis and Recognition),19(2),119-138。网址:https://link.springer.com/article/10.1007/s10032-016-0259-1。 [3] Shigarov A., Paramonov V., Belykh P., Bondarev A. (2016). 电子表格任意表格的基于规则规范化处理. 第22届国际信息与软件技术会议论文集(Proc. 22nd Int. Conf. on Information and Software Technologies),第78-91页。网址:http://link.springer.com/chapter/10.1007/978-3-319-46254-7_7。

创建时间:
2017-06-25
二维码
社区交流群
二维码
科研交流群
商业服务