LettuceDetect benchmark
收藏资源简介:
该数据集是由KR实验室等机构构建的统一基准,旨在评估检索增强生成(RAG)系统在结构化输入下的幻觉检测能力。数据集包含74,285个新构建的示例,涵盖代码、工具输出、结构化文档及自然语言RAG数据,具体来源包括SWE-bench代码、Squeez工具输出、ACL论文块、README文档和维基百科标记文本。数据集的创建过程采用基于编辑的标注方法,从正确的接地答案出发,注入局部幻觉并精确标注字符偏移,以确保标签质量。该数据集主要应用于代码代理、开发助手及文档系统等领域,旨在解决结构化上下文(如源代码、工具输出)中生成答案的幻觉检测问题,提升生成系统的可靠性和准确性。
This dataset is a unified benchmark constructed by institutions including KR Labs, aiming to evaluate the hallucination detection performance of retrieval-augmented generation (RAG) systems on structured inputs. It contains 74,285 newly constructed examples covering code, tool outputs, structured documents and natural language RAG data, with specific sources encompassing SWE-bench code, Squeez tool outputs, ACL paper chunks, README documents and Wikipedia tagged text. The dataset was developed using an edit-based annotation workflow: starting from correctly grounded reference answers, local hallucinations are injected, and character offsets are precisely labeled to ensure the quality of the annotation labels. This dataset is primarily deployed in scenarios including code agents, development assistants and documentation systems, with the objective of addressing hallucination detection issues for generated outputs within structured contexts (e.g., source code, tool outputs), and enhancing the reliability and accuracy of generative systems.

- 1Beyond Document Grounding: Span-Level Hallucination Detection over Code, Tool Output, and DocumentsKR实验室; MBZUAI; 麦吉尔大学; 维也纳工业大学 · 2026年



