Qualcomm Interactive Video Dataset (IVD)
收藏资源简介:
高通交互式视频数据集(IVD)是一个专门为评估AI模型在实时情境下视觉理解能力而设计的多模态数据集。该数据集由2900个视频剪辑组成,每个视频都配有与视频内容同步的问题和答案对。这些问题和答案经过人工标注,并且包含了回答问题的时间戳。数据集的视频展示了各种不同的场景、行为和对象,旨在训练和评估AI系统在理解视觉场景方面的能力。IVD数据集能够为在线大型多模态模型的情境音频-视觉推理研究提供支持,并可用于构建能够实时与用户互动的对话系统。
Qualcomm Interactive Video Dataset (IVD) is a multimodal dataset specifically designed to evaluate the visual understanding capabilities of AI models in real-time scenarios. This dataset comprises 2900 video clips, each paired with question-answer pairs synchronized with the corresponding video content. All question-answer pairs are manually annotated and equipped with timestamps marking the temporal positions relevant to answering the questions. The videos in the dataset cover a diverse range of scenarios, behaviors and objects, and are intended to train and evaluate AI systems' visual scene comprehension abilities. The IVD dataset can provide support for research on situational audio-visual reasoning of large online multimodal models, and can also be used to build conversational systems capable of real-time interaction with users.




