KRLabsOrg/lettucedetect-code-hallucination
收藏资源简介:
该数据集名为LettuceDetect Grounded Hallucination Dataset,是一个用于幻觉检测的数据集,包含对基于结构化上下文的LLM响应的token级幻觉标注。数据来源于五个方面:源代码、开发者工具输出、学术论文、GitHub READMEs和Wikipedia。这是LettuceDetect数据收集的一部分。每个样本将一个基于上下文的LLM答案与正确或包含最小扰动、字符跨度标注的幻觉配对。所有跨度使用统一的分类法,因此来源共享一个标签空间,可以通过dataset/context_modality字段联合训练或过滤分开。
Token-level hallucination annotations on LLM responses grounded in structured context across five sources — source code, developer-tool output, academic papers, GitHub READMEs, and Wikipedia. Part of the LettuceDetect data collection. Every sample pairs a grounded context with an LLM answer that is either correct or contains a minimally perturbed, character-span-annotated hallucination. All spans use one unified taxonomy, so the sources share a single label space and can be trained jointly or filtered apart via the `dataset` / `context_modality` fields.




