LFED
收藏资源简介:
LFED是一个专为评估大型语言模型在长篇文学小说理解和推理能力上的数据集。该数据集由天津大学智能与计算学部创建,包含95部中文原著或翻译的文学小说,涵盖多个世纪的广泛主题。数据集通过众包方式构建,严格控制质量,最终形成1304个问题,涉及8种不同的问题类型,旨在全面评估语言模型在事实理解、逻辑推理、上下文理解、常识推理和价值判断等方面的能力。LFED的应用领域主要集中在评估和提升大型语言模型在文学领域的理解和推理能力,以解决现有数据集无法充分评估大型模型的问题。
LFED is a specialized dataset developed for evaluating the long-form literary fiction comprehension and reasoning capabilities of large language models (LLMs). Created by the Faculty of Intelligence and Computing, Tianjin University, the dataset includes 95 Chinese original or translated literary novels covering a wide range of themes spanning multiple centuries. Constructed through crowdsourcing with rigorous quality control, it ultimately consists of 1304 questions across 8 distinct question types, aiming to comprehensively assess language models' abilities in factual understanding, logical reasoning, contextual comprehension, commonsense reasoning, value judgment, and other relevant aspects. The primary application fields of LFED focus on evaluating and enhancing the literary comprehension and reasoning capabilities of large language models, so as to address the limitation that existing datasets cannot adequately evaluate large-scale models.

- 1LFED: A Literary Fiction Evaluation Dataset for Large Language Models天津大学智能与计算学部 · 2024年



