遇见数据集

tsinghua-ee/video-SALMONN_2_testset

收藏
Hugging Face2025-08-20 更新2025-08-09 收录
官方服务:

资源简介:

video-SALMONN 2 Benchmark数据集是一个视频到文本的任务,它要求生成与视频和音频相对应的字幕。数据集以Apache-2.0许可发布,支持英语语言。使用该数据集时,用户需要按照指定的格式组织生成的字幕,并通过提供的脚本进行评估。

The video-SALMONN 2 Benchmark dataset is a video-to-text task that requires generating captions corresponding to the video and audio. The dataset is released under the Apache-2.0 license and supports the English language. When using this dataset, users need to organize the generated captions in a specified format and evaluate them using the provided script.

提供机构:
tsinghua-ee
搜集汇总
数据集介绍
tsinghua-ee/video-SALMONN_2_testset 数据集图片
背景与挑战
背景概述
该数据集是video-SALMONN 2模型的测试集,用于视频文本到文本任务,专注于英语视频内容,基于Apache 2.0许可证发布。数据集旨在评估音频-视觉大语言模型生成视频标题的能力,但当前存在数据列不匹配问题,导致无法正常预览和使用。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务