HiEviDR-Bench
收藏资源简介:
HiEviDR-Bench是一个用于评估多模态深度研究中层次证据聚合的基准测试。它旨在评估模型是否能够正确地从大规模异构来源中检索、连接和综合证据,而不仅仅是生成流畅的最终答案或报告。该基准测试提供了对中间证据聚合过程的明确监督,每个实例都标注了一个证据图,捕捉证据如何被选择、跨源链接并聚合为中间主张和最终结论。发布的基准测试包含3,407个研究导向的问题、支持性语料数据、层次证据聚合的证据图、纯文本和多模态设置,以及开放领域和学术领域子集。HiEviDR-Bench支持细粒度分析,将每个示例制定为层次证据聚合问题,并开发了一个面向可追溯性的评估框架,包含五个维度:报告质量、证据可追溯性、引用准确性、主张验证和答案正确性。数据集结构包括问题ID、问题文本、参考答案、多模态输入和输出、证据ID列表、证据图和证据项等字段。
HiEviDR-Bench is a benchmark for evaluating hierarchical evidence aggregation in multimodal in-depth research. It aims to assess whether models can correctly retrieve, connect, and synthesize evidence from large-scale heterogeneous sources, rather than merely generating fluent final answers or reports. This benchmark provides explicit supervision over the intermediate evidence aggregation process, with each instance annotated with an evidence graph that captures how evidence is selected, linked across sources, and aggregated into intermediate claims and final conclusions. The released benchmark contains 3,407 research-oriented questions, supporting corpora, evidence graphs for hierarchical evidence aggregation, both plain-text and multimodal settings, as well as open-domain and academic domain subsets. HiEviDR-Bench supports fine-grained analysis by framing each example as a hierarchical evidence aggregation problem, and has developed a traceability-oriented evaluation framework covering five dimensions: report quality, evidence traceability, citation accuracy, claim verification, and answer correctness. The dataset structure includes fields such as question ID, question text, reference answer, multimodal input and output, evidence ID list, evidence graph, and evidence items.
HiEviDR-Bench数据集概述
数据集基本信息
- 名称: HiEviDR-Bench
- 许可证: Apache-2.0
- 任务类别: 视觉问答
- 主要语言: 英语
核心目标
HiEviDR-Bench是一个用于评估多模态深度研究中分层证据聚合的基准。它旨在评估模型是否能从大规模异构来源中正确检索、连接和综合证据,而不仅仅是生成流畅的最终答案或报告。
数据集构成
- 研究导向问题: 3,407个
- 支持语料库数据: 包含
- 证据图: 用于分层证据聚合
- 模态设置: 纯文本和多模态
- 领域子集: 开放领域和学术领域
任务描述
给定一个研究导向的问题,系统需要从多模态语料库中检索并聚合相关证据,然后生成:
- 结构化或长形式的报告
- 有依据的答案
涵盖模态
- 文本
- 多模态
涵盖领域
- 维基百科
- arXiv
数据结构
典型数据样本包含以下字段:
question_id: 问题实例的唯一标识符question: 研究导向的问题answer: 问题的参考答案mm_inputs: 与问题相关的多模态输入mm_outputs: 与参考答案相关的多模态输出evidence_ids: 与此问题相关的证据项ID列表evidence_graph: 描述证据如何支持中间主张和最终结论的分层证据图evidence_items: 原始证据项的字典ret2cid: 检索结果与引用/证据标识符之间的可选映射
评估框架
采用面向可追溯性的评估框架,包含五个维度:
- 报告质量
- 证据可追溯性
- 引用准确性
- 主张验证
- 答案正确性
其他信息
- 项目页面: https://boggysyb.github.io/HiEviDR-Bench.github.io/
- 联系方式: syb2000417@stu.pku.edu.cn
- 初始发布日期: 2026-04-17




