lucky-lance/OmniInteract
收藏资源简介:
OmniInteract是一个针对实时全模态大语言模型的流式基准测试,通过其本地在线推理在连续音视频流上进行评估。用户查询和环境声音位于音频轨道,视觉事件位于视频中,模型必须决定是否、何时以及如何响应——无需预知未来内容。该数据集包含250个视频和1,430个时间锚定的响应槽位,分为两个任务:1Q1A(210个视频/1,062个槽位)用于局部单响应交互(实时、主动和嵌套),1QnA(40个视频/368个槽位)用于长时程连续任务监控(一个指令→多个时间锚定答案)。领域涵盖中文日常生活交互(家庭、健身房、博物馆、购物等)和英文数学推理。
OmniInteract is a streaming benchmark for real-time omnimodal LLMs, evaluated through their native online inference over continuous audio-visual streams. User queries and ambient sounds live in the audio track, visual events live in the video, and a model must decide whether, when, and what to respond — without lookahead to future content. The dataset includes 250 videos and 1,430 temporally grounded response slots, with two tasks: 1Q1A (210 videos / 1,062 slots) for localized single-response interaction (real-time, proactive, and nested), and 1QnA (40 videos / 368 slots) for long-horizon continuous task monitoring (one instruction → many time-grounded answers). Domains cover Chinese daily-life interaction (home, gym, museum, shopping, …) and English mathematical reasoning.





