FactIR
收藏资源简介:
FactIR数据集是由Factiverse公司生产日志中的证据检索数据及人工标注信息构成的,旨在为事实核查提供一个真实的零样本开放域检索基准。该数据集包含1413个声明确认-证据对的相关性标注,以及90047篇文档的语料库集合。数据集中的声明是有机生成的,由用户在涉及政治、健康、经济等多个主题的任务中产生,并通过专业的事实核查者和具有新闻背景的个人进行验证。该数据集的构建旨在填补事实核查研究中缺乏真实世界开放域检索基准的空白。
The FactIR dataset is constructed from evidence retrieval data and manually annotated information sourced from the production logs of Factiverse, serving as a realistic zero-shot open-domain retrieval benchmark for fact-checking. This dataset includes 1413 relevance annotations for claim-evidence pairs, alongside a corpus of 90,047 documents. The claims within the dataset are organically generated by users across tasks covering diverse topics such as politics, health, and economics, and validated by professional fact-checkers and individuals with journalistic backgrounds. The construction of this dataset aims to fill the gap in real-world open-domain retrieval benchmarks for fact-checking research.

- 1FactIR: A Real-World Zero-shot Open-Domain Retrieval Benchmark for Fact-Checking荷兰代尔夫特理工大学 · 2025年



