PresentEval
收藏资源简介:
PresentEval是一个多模态演示评估基准,源自论文《PresentAgent-2: Towards Generalist Multimodal Presentation Agents》。该基准旨在评估能够将开放式用户查询转化为带旁白演示视频的智能体框架,核心目标是衡量智能体在研究主题、检索多模态资源以及跨不同交付模式传递结构化内容方面的综合能力。评估涵盖三种具体模式:1) 单讲者演示:生成单讲者旁白的演示视频;2) 讨论:创建具有结构化角色的多讲者演示,角色包括提问引导、解释概念、澄清细节和总结要点;3) 互动:评估基于生成的幻灯片、脚本、检索到的证据和演示上下文来回答观众问题的能力。评估采用客观测验评估(使用视觉语言模型作为观众,根据生成的视频和音频转录回答多项选择题以衡量知识传递效果)和主观评分(使用视觉语言模型评委根据内容质量、媒体相关性、对话自然度和互动依据等模式特定标准进行1-5分打分)。该数据集适用于文本到视频生成、多模态智能体评估以及自动演示生成等任务场景。
PresentEval is a multimodal demonstration evaluation benchmark derived from the paper *PresentAgent-2: Towards Generalist Multimodal Presentation Agents*. This benchmark is designed to evaluate agent frameworks that can convert open-ended user queries into presentation videos with voiceover, with the core objective of assessing the comprehensive capabilities of AI agents in selecting appropriate research topics, retrieving multimodal resources, and delivering structured content across diverse delivery modes. The evaluation covers three specific modes: 1) Single-speaker Presentation: generating presentation videos with voiceover from a single speaker; 2) Discussion Session: creating multi-speaker presentations with structured roles, including roles for guiding inquiries, explaining concepts, clarifying details, and summarizing key takeaways; 3) Interaction: evaluating the ability to answer audience questions based on generated slides, scripts, retrieved evidence, and presentation context. The evaluation employs two assessment paradigms: objective quiz evaluation (using vision-language models as simulated audiences, which answer multiple-choice questions based on the generated videos and audio transcripts to measure the effectiveness of knowledge transfer) and subjective scoring (using vision-language model judges to assign scores ranging from 1 to 5 according to mode-specific criteria including content quality, media relevance, naturalness of dialogue, and grounding of interactions). This dataset is applicable to task scenarios such as text-to-video generation, multimodal agent evaluation, and automatic presentation generation.





