InfoTabS
收藏资源简介:
InfoTabS 包含基于前提的人类书面文本假设,这些前提是从维基百科信息框中提取的表格。在本文中,我们观察到半结构化表格文本无处不在;理解它们不仅需要理解文本片段的含义,还需要理解它们之间的隐含关系。我们认为,这些数据可以证明是理解我们如何推理信息的试验场。为了研究这一点,我们引入了一个名为 INFOTABS 的新数据集,其中包括基于前提的人工编写的文本假设,这些前提是从维基百科信息框中提取的表格。我们的分析表明,前提的半结构化、多领域和异构性质允许进行复杂的、多方面的推理。实验表明,虽然人类注释者同意表格-假设对之间的关系,但一些标准的建模策略在该任务中并不成功,这表明关于表格的推理可能会带来困难的建模挑战。
InfoTabS consists of human-written textual hypotheses grounded on premises, where the premises are tables extracted from Wikipedia infoboxes. In this work, we observe that semi-structured tabular text is pervasive; understanding such text requires not only comprehending the meaning of individual textual fragments, but also the implicit relationships between them. We posit that such data can serve as a valuable testbed for understanding how we reason about information. To investigate this, we introduce a novel dataset named INFOTABS, which comprises human-written textual hypotheses grounded on premises that are tables extracted from Wikipedia infoboxes. Our analysis reveals that the semi-structured, multi-domain, and heterogeneous nature of the premises enables complex, multi-faceted reasoning. Experiments demonstrate that while human annotators reach consensus on the relationships between table-hypothesis pairs, several standard modeling strategies perform poorly on this task, indicating that reasoning over tabular data poses challenging modeling hurdles.




