DetectBench
收藏资源简介:
DetectBench是一个专为评估大型语言模型在复杂信息中识别关键信息和进行多步推理能力而设计的阅读理解数据集。该数据集包含3928个问题,每个问题伴随一个平均长度为190个令牌的段落。DetectBench的特点包括:关键信息不直接出现在上下文中,需要结合上下文中的多个线索推导出更关键的线索,且上下文中包含大量误导性和无关信息。数据集从开源平台收集大量侦探谜题,并重写为包含上下文、问题、选项、答案及答案解释的格式。DetectBench旨在通过模拟侦探谜题中的复杂故事、情境和角色互动,挑战模型在检测和推理线索方面的能力,以解决实际问题。
DetectBench is a reading comprehension dataset specifically designed to evaluate the capabilities of large language models (LLMs) in identifying critical information from complex contexts and conducting multi-step reasoning. This dataset comprises 3,928 questions, each paired with a passage averaging 190 tokens in length. The distinguishing features of DetectBench are as follows: critical information does not directly appear in the provided context; multiple clues within the context need to be integrated to deduce more critical clues; and the context contains a substantial amount of misleading and irrelevant information. The dataset is compiled from a large number of detective puzzles sourced from open-source platforms, then rewritten into a standardized format that includes context, questions, options, correct answers, and answer explanations. DetectBench aims to challenge models' abilities in clue detection and reasoning by simulating complex narratives, scenarios, and character interactions in detective puzzles, so as to address practical real-world problems.




