TabFact
收藏资源简介:
TabFact是一个大规模的数据集,用于基于表格的事实验证。该数据集由加州大学圣巴巴拉分校和腾讯AI Lab创建,包含16,000个维基百科表格作为证据,用于验证118,000个人工标注的自然语言陈述,这些陈述被标记为'ENTAILED'或'REFUTED'。TabFact的挑战在于它涉及软语言推理和硬符号推理。数据集创建过程中,通过众包方式收集了不同难度级别的陈述,并进行了质量控制以确保数据的准确性。TabFact的应用领域包括自然语言理解和语义表示的研究,旨在解决基于结构化证据的事实验证问题。
TabFact is a large-scale dataset for table-based fact verification. Created by the University of California, Santa Barbara and Tencent AI Lab, it contains 16,000 Wikipedia tables as evidence, which are used to validate 118,000 manually annotated natural language statements labeled as 'ENTAILED' or 'REFUTED'. The core challenge of TabFact lies in its integration of both soft linguistic reasoning and hard symbolic reasoning. During the dataset construction process, statements of varying difficulty levels were collected via crowdsourcing, and quality control measures were adopted to ensure the accuracy of the dataset. The application domains of TabFact include research in natural language understanding and semantic representation, with the goal of solving structured evidence-based fact verification problems.




