A subset of the Scientific Literature Comparison Tables Dataset
收藏资源简介:
The Scientific Literature Comparison Table (SLCT) dataset was collected using the arXiv and Semantic Scholar APIs. It underwent a series of processing steps, including manual inspection and editing. The processing steps are summarized as follows: 1) Downloading Survey Papers’ LaTeX files using the Arxiv API. 2) Preprocessing LaTeX files to HTML format. 3) Extracting tables from the HTML files. 4) Creating a Golden Table as a reference. 5) Generating descriptions for column headers. 6) Acquiring citation data. 7) Finalizing the dataset.
科学文献对比表(Scientific Literature Comparison Table,SLCT)数据集通过arXiv与Semantic Scholar的应用程序编程接口(Application Programming Interface,API)采集而来,其历经一系列包含人工核验与编辑的处理流程。具体处理步骤总结如下: 1. 依托arXiv API下载综述论文的LaTeX源文件; 2. 将LaTeX源文件预处理为HTML格式; 3. 从HTML文件中提取表格; 4. 构建作为参考基准的金标准表格(Golden Table); 5. 为表格列标题生成描述文本; 6. 获取引用数据; 7. 完成数据集的最终定稿。



