FRESHQAPARALLEL, SEAREFUSE, TRUEFALSEMULTILANG
收藏资源简介:
本研究构建了一个多语言评估套件,包括三种类型的知识边界数据:具有真或假前提的问题、以实体为中心的可答/不可答问题以及真假陈述。数据集包括FRESHQAPARALLEL(扩展自FRESHQA,包含平行真/假前提问题)、SEAREFUSE(包含关于不存在实体的不可答问题和可答问题)和TRUEFALSEMULTILANG(将TRUEFALSE数据集翻译成多种语言)。这些数据集旨在分析大规模语言模型如何跨语言泛化其知识边界的认知。
This study constructs a multilingual evaluation suite encompassing three types of knowledge boundary data: questions with true or false premises, entity-centered answerable/unanswerable questions, and true/false statements. The suite includes three core datasets: FRESHQAPARALLEL (extended from FRESHQA, containing parallel pairs of true/false premise questions), SEAREFUSE (comprising answerable and unanswerable questions about non-existent entities), and TRUEFALSEMULTILANG (translated from the original TRUEFALSE dataset into multiple languages). This evaluation suite is designed to analyze how large language models (LLMs) generalize their awareness of knowledge boundaries across languages.




