VISTA
收藏资源简介:
VISTA数据集是由Saarland University、University of Cambridge和University of Edinburgh共同创建的一个英语多模态数据集,专为科学领域的视频到文本摘要任务设计。该数据集包含来自计算语言学和机器学习领域领先会议的18599对视频和相应论文摘要的配对,涵盖了从2020年到2024年的内容。数据集的平均视频长度为6.76分钟,摘要平均包含192.62个token,体现了其语料库的多样性和复杂性。该数据集的创建旨在解决科学视频摘要领域的挑战,并为视频到文本摘要任务提供基准。
The VISTA dataset is an English multimodal dataset jointly created by Saarland University, University of Cambridge and University of Edinburgh, specifically tailored for video-to-text summarization tasks in the scientific domain. This dataset includes 18,599 pairs of videos and their corresponding paper abstracts sourced from top conferences in computational linguistics and machine learning, covering content from 2020 to 2024. The average duration of videos in the dataset is 6.76 minutes, and the average number of tokens per abstract is 192.62, reflecting the diversity and complexity of this corpus. The dataset was developed to address the challenges in the field of scientific video summarization and provide a benchmark for video-to-text summarization tasks.




