FEVER (Fact Extraction and VERification)
收藏资源简介:
FEVER 是一个公开可用的数据集,用于对文本源进行事实提取和验证。 它由 185,445 个声明组成,根据维基百科页面的介绍部分手动验证并分类为 SUPPORTED、REFUTED 或 NOTENOUGHINFO。对于前两个类别,系统和注释者还需要返回句子组合,形成支持或反驳该主张的必要证据。 这些声明是由人类注释者从维基百科中提取声明并以各种方式对其进行变异产生的,其中一些是改变含义的。每个声明的验证是由注释者在单独的注释过程中进行的,他们知道页面而不是提取原始声明的句子,因此在 31.75% 的声明中,超过一个句子被认为是适当的证据。索赔要求在 16.82% 的案件中组合来自多个句子的证据。此外,在 12.15% 的索赔中,该证据取自多页。
FEVER is a publicly available dataset for fact extraction and verification over textual sources. It consists of 185,445 claims, which are manually verified and categorized into SUPPORTED, REFUTED, or NOTENOUGHINFO based on the introduction sections of Wikipedia pages. For the first two categories, systems and annotators are also required to return sentence combinations that form the necessary evidence to support or refute the claim. These claims are generated by human annotators extracting statements from Wikipedia and mutating them in various ways, some of which alter the meaning. The verification of each claim is conducted by annotators in a separate annotation process, who are provided with the relevant pages rather than the exact sentences from which the original claim was extracted. As a result, in 31.75% of claims, more than one sentence is considered appropriate evidence. Claims require combining evidence from multiple sentences in 16.82% of cases. Additionally, in 12.15% of claims, the evidence is taken from multiple pages.

- FEVER数据集首次发布,旨在为事实验证任务提供一个标准化的基准。
- FEVER共享任务启动,吸引了全球多个研究团队参与,推动了事实验证技术的发展。
- FEVER数据集扩展,增加了更多的标注数据和新的验证任务,进一步丰富了数据集的内容和应用场景。
- 1FEVER: A Large-scale Dataset for Fact Extraction and VERificationUniversity of Oxford, University of Sheffield, University of Lisbon · 2018年
- 2Fact or Fiction: Verifying Scientific ClaimsUniversity of Washington, Allen Institute for AI · 2020年
- 3Fact Checking in Community ForumsUniversity of California, Berkeley · 2021年
- 4Fact-Checking with Insufficient EvidenceUniversity of Amsterdam, University of Copenhagen · 2022年
- 5FEVEROUS: Fact Extraction and VERification Over Unstructured and Structured informationUniversity of Lisbon, University of Sheffield · 2021年



