遇见数据集

AIGeeksGroup/PresentEval

收藏
Hugging Face2026-05-13 更新2026-06-14 收录
官方服务:

资源简介:

PresentEval是一个多模态演示基准数据集,旨在评估代理框架将开放式用户查询转化为带旁白的演示视频的能力。该数据集衡量代理在研究主题、检索多模态资源以及在三种不同交付模式下提供结构化内容的能力:单人演示(生成单人旁白的演示视频)、讨论(创建具有结构化角色的多发言人演示,包括提出引导性问题、解释概念、澄清细节和总结要点)以及互动(评估基于生成的幻灯片、脚本、检索的证据和演示上下文回答观众问题的能力)。评估方法采用客观测验评估(通过视觉语言模型作为观众回答多项选择题)和主观评分(使用视觉语言模型评委根据内容质量、媒体相关性、对话自然度和互动基础等模式特定标准进行1-5分评分)。

PresentEval is a multimodal presentation benchmark introduced in the paper PresentAgent-2: Towards Generalist Multimodal Presentation Agents. The benchmark is designed to evaluate agentic frameworks that transform open-ended user queries into narrated presentation videos. It measures an agents ability to research topics, retrieve multimodal resources, and deliver structured content across three distinct delivery modes: Single Presentation (generates a single-speaker narrated presentation video), Discussion (creates a multi-speaker presentation with structured roles for asking guiding questions, explaining concepts, clarifying details, and summarizing key points), and Interaction (evaluates the ability to answer audience questions grounded in generated slides, scripts, retrieved evidence, and presentation context). Evaluation employs objective quiz evaluation (a VLM acts as an audience member to answer multiple-choice questions) and subjective scoring (uses a VLM judge to assign 1–5 scores based on mode-specific criteria such as content quality, media relevance, dialogue naturalness, and interaction grounding).

提供机构:
AIGeeksGroup
二维码
社区交流群
二维码
科研交流群
商业服务