Paper2Video
收藏资源简介:
Paper2Video是一个包含101篇论文与其对应的作者录制演讲视频、幻灯片和演讲者元数据的基准数据集。数据集涵盖了机器学习、计算机视觉和自然语言处理等多个领域,每篇论文平均包含13.3K个词(3.3K个tokens)、44.7个图表和28.7页内容,提供了多模态长文档输入。数据集旨在解决学术演示视频生成中的多模态长上下文理解、多轮代理任务、个性化演讲者合成和时空定位等挑战。
Paper2Video is a benchmark dataset consisting of 101 academic papers paired with their corresponding author-recorded presentation videos, slides, and speaker metadata. The dataset covers multiple research fields including machine learning, computer vision, and natural language processing. On average, each paper contains 13.3K words (3.3K tokens), 44.7 figures and tables, and 28.7 pages of content, providing multi-modal long-document inputs. This dataset is designed to address core challenges in academic presentation video generation, such as multi-modal long-context understanding, multi-turn agent tasks, personalized speaker synthesis, and spatiotemporal localization.

- 1通过新加坡国立大学 · 2025年



