HalluMix Benchmark
收藏资源简介:
HalluMix Benchmark是一个大规模、领域多样化的数据集,专门设计用于评估真实生成场景中的幻觉检测。该数据集包括来自多个任务的示例,包括摘要、问答和自然语言推理,并涵盖广泛的领域,如医疗保健、法律、科学和新闻。每个实例都包含一个多文档上下文和一个响应,以及二进制幻觉标签,指示响应是否忠实于提供的文档。数据集旨在反映现实世界的信息检索场景,并评估现有幻觉检测方法在不同任务、文档长度和输入表示上的性能差异。
HalluMix Benchmark is a large-scale, domain-diverse dataset specifically designed to evaluate hallucination detection in real-world generative scenarios. This dataset includes examples from multiple tasks such as summarization, question answering, and natural language inference, covering a wide range of domains including healthcare, law, science, and journalism. Each instance contains a multi-document context, a response, and a binary hallucination label indicating whether the response is faithful to the provided documents. The dataset is intended to reflect real-world information retrieval scenarios and assess the performance differences of existing hallucination detection methods across different tasks, document lengths, and input representations.




