遇见数据集

llm-semantic-router/halueval-spans

收藏
Hugging Face2026-01-09 更新2026-02-07 收录
官方服务:

资源简介:

HaluEval Span-Level Dataset (LLM-Detected) 是一个高质量的跨度级幻觉检测数据集,通过使用 Qwen2.5-72B-Instruct 从 HaluEval 转换而来,提供精确的跨度检测和 RAGTruth 兼容的标签。数据集包含 10,000 个摘要样本,其中 8,905 个样本检测到幻觉(89%),总共有 16,359 个幻觉跨度。数据集包含四种 RAGTruth 兼容的标签类型:明显无根据信息、明显冲突、微妙无根据信息和微妙冲突。该数据集专为跨度级幻觉检测任务设计,如训练令牌级幻觉检测器、多标签分类和研究细粒度幻觉类型。它与 RAGTruth 兼容,适合联合训练和评估。

The HaluEval Span-Level Dataset (LLM-Detected) is a high-quality span-level hallucination detection dataset converted from HaluEval using Qwen2.5-72B-Instruct for precise span detection and RAGTruth-compatible labeling. It contains 10,000 summarization samples with 8,905 samples containing detected hallucinations (89%) and 16,359 total hallucinated spans. The dataset features four RAGTruth-compatible label types: Evident Baseless Info, Evident Conflict, Subtle Baseless Info, and Subtle Conflict. Designed for span-level hallucination detection tasks, it is suitable for training token-level hallucination detectors, multi-label classification, and research on fine-grained hallucination types. It is compatible with RAGTruth and suitable for combined training and evaluation.

提供机构:
llm-semantic-router
搜集汇总
数据集介绍
llm-semantic-router/halueval-spans 数据集图片
背景与挑战
背景概述
该数据集是一个高质量的跨度级幻觉检测数据集,由HaluEval通过Qwen2.5-72B-Instruct转换而来,包含10,000个摘要样本,其中89%检测到幻觉,总计16,359个幻觉跨度,并提供四种RAGTruth兼容的标签类型。它专为跨度级幻觉检测任务设计,适用于训练检测器、多标签分类和细粒度研究,且与RAGTruth框架兼容,支持联合训练和评估。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务