DTVLT
收藏资源简介:
DTVLT是由中国科学院自动化研究所等机构创建的多模态视觉语言跟踪基准,旨在通过多样化的文本描述提升视频理解算法的性能。数据集包含13134条视频数据,覆盖了短期跟踪、长期跟踪和全局实例跟踪三个主流任务。创建过程中,利用大型语言模型(LLM)生成多粒度的文本描述,以丰富视频内容的语义信息。DTVLT的应用领域主要集中在视觉语言跟踪和视频理解,旨在解决传统单模态跟踪算法在复杂视频内容理解中的局限性。
DTVLT is a multimodal vision-language tracking benchmark developed by institutions including the Institute of Automation, Chinese Academy of Sciences, aiming to improve the performance of video understanding algorithms via diverse textual descriptions. The dataset contains 13,134 video samples, covering three mainstream tasks: short-term tracking, long-term tracking, and global instance tracking. During its construction, large language models (LLMs) were utilized to generate multi-granularity textual descriptions to enrich the semantic information of video content. The application domains of DTVLT mainly focus on vision-language tracking and video understanding, and it is designed to address the limitations of traditional unimodal tracking algorithms in complex video content understanding.




