CATS
收藏资源简介:
CATS是一个大规模、高质量的中文答案到序列数据集,由中国科学院信息工程研究所和阿里巴巴达摩院共同创建。该数据集旨在为实际的TableQA系统生成文本描述,包含43,369个答案到序列的例子,远超现有数据集。CATS通过手动标注所有收集的SQL-表对来确保数据质量,并采用统一图转换方法来弥合输入SQL和表之间的结构差距,将任务转换为图到文本问题。此外,CATS还引入了节点段嵌入以更好地保留原始结构信息。该数据集的应用领域主要集中在提高TableQA系统的用户友好性和交互性,解决现有数据集在语言多样性和实际应用场景方面的不足。
CATS is a large-scale, high-quality Chinese answer-to-sequence dataset co-developed by the Institute of Information Engineering, Chinese Academy of Sciences and Alibaba DAMO Academy. This dataset aims to generate textual descriptions for real-world TableQA systems, containing 43,369 answer-to-sequence examples, which is significantly larger than existing datasets. CATS ensures data quality by manually annotating all collected SQL-table pairs, and adopts a unified graph transformation method to bridge the structural gap between input SQL and tables, thus converting the task into a graph-to-text problem. Additionally, CATS introduces node segment embeddings to better preserve the original structural information. The application scenarios of this dataset mainly focus on improving the user-friendliness and interactivity of TableQA systems, addressing the shortcomings of existing datasets in terms of linguistic diversity and real-world application scenarios.




