Poly-FEVER
收藏资源简介:
Poly-FEVER是一个大规模的多语言事实验证基准数据集,由美国北德克萨斯大学的研究团队创建。该数据集包含11种语言的77,973条标注事实主张,来源于FEVER、Climate-FEVER和SciFact。Poly-FEVER旨在评估大型语言模型中虚假信息的检测,特别关注跨语言的一致性。数据集覆盖了艺术、音乐、科学、生物学和 history 等多个主题,支持跨语言的事实验证研究,推动了对大型语言模型中虚假信息模式的深入理解。
Poly-FEVER is a large-scale multilingual fact verification benchmark dataset created by a research team from the University of North Texas, USA. This dataset contains 77,973 labeled factual claims across 11 languages, sourced from FEVER, Climate-FEVER, and SciFact. Poly-FEVER aims to evaluate misinformation detection in large language models, with particular focus on cross-lingual consistency. The dataset covers multiple topics including art, music, science, biology, and history, supporting cross-lingual fact verification research and advancing in-depth understanding of misinformation patterns in large language models.
Poly-FEVER数据集概述
数据集基本信息
- 名称: Poly-FEVER
- 语言: 英语(en)、中文(zh)、印地语(hi)、阿拉伯语(ar)、孟加拉语(bn)、日语(ja)、韩语(ko)、泰米尔语(ta)、泰语(th)、格鲁吉亚语(ka)、阿姆哈拉语(am)
- 数据规模: 10K<n<100K
- 任务类型: 文本分类
数据集描述
Poly-FEVER是一个多语言事实验证基准数据集,旨在评估大型语言模型(LLMs)中的幻觉检测能力。该数据集通过将声明翻译成11种语言,扩展了三个广泛使用的事实核查数据集:FEVER、Climate-FEVER和SciFact。
关键特征
- 包含77,973个事实声明
- 二元标签(SUPPORTS或REFUTES)
- 覆盖多个领域:艺术、科学、政治和历史
- 资助方: Google Cloud Translation
数据来源
- FEVER: https://fever.ai/resources.html
- CLIMATE-FEVER: https://www.sustainablefinance.uzh.ch/en/research/climate-fever.html
- SciFact: https://huggingface.co/datasets/allenai/scifact
相关论文
- 论文链接: https://huggingface.co/papers/2503.16541
数据集创建信息
原始数据集
- FEVER
- Climate-FEVER
- SciFact
注意事项
- 用户应注意数据集可能存在的风险、偏见和限制
- 更多详细信息待补充

- 1Poly-FEVER: A Multilingual Fact Verification Benchmark for Hallucination Detection in Large Language Models美国北德克萨斯大学 · 2025年



