noahschaffer/clasp-audioset
收藏官方服务:
资源简介:
该数据集包含图像及其对应的文本描述,每个样本还包括标签列表、YouTube视频ID和起始时间戳。数据集分为训练集(17,372个样本)和评估集(15,778个样本),图像和描述用于多模态学习任务。
This dataset contains images paired with textual captions, along with a list of labels, YouTube video ID, and start timestamp. It is split into training (17,372 examples) and evaluation (15,778 examples) sets, intended for multimodal learning tasks.
提供机构:
noahschaffer



