TriSense-2M
收藏资源简介:
TriSense-2M是一个大规模的多模态数据集,包含超过200万条注释。每个视频实例都包括在视觉、音频和语音模态上基于事件进行注释,并且具有灵活的组合和模态的自然缺失。数据集支持各种场景,并包括平均时长为905秒的长视频,这显著长于现有数据集中的视频,从而能够实现更深层次和更真实的时序理解。重要的是,查询使用高质量的母语语言,与时间注释对齐,并且跨越不同的模态配置,以促进鲁棒的多模态学习。
TriSense-2M is a large-scale multimodal dataset with over 2 million annotated instances. Each video instance is annotated at the event level across visual, audio, and speech modalities, supporting flexible modality combinations and natural modality absence. The dataset covers diverse scenarios and includes long videos with an average duration of 905 seconds, which is significantly longer than videos in existing datasets, enabling deeper and more realistic temporal understanding. Crucially, the queries are formulated in high-quality native languages, aligned with temporal annotations, and cover diverse modality configurations to facilitate robust multimodal learning.

- 1Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM浙江大学 · 2025年



