SIV-Bench
收藏资源简介:
SIV-Bench是一个视频基准数据集,用于严格评估多模态大型语言模型(MLLMs)在社交场景理解(SSU)、社交状态推理(SSR)和社交动态预测(SDP)方面的能力。该数据集包含2792个视频片段和8792个精心生成的问答对,源自于人类与LLM协作的管道。它最初是从TikTok和YouTube收集的,涵盖广泛的视频类型、展示风格和语言文化背景。SIV-Bench还专门设置了分析不同文本线索影响的环境,包括原始屏幕文本、添加的对话或无文本。在领先的MLLMs上进行的综合实验表明,模型在处理SSU方面表现得很好,但在SSR和SDP方面却存在显著困难,其中关系推理(RI)是一个严重的瓶颈。SIV-Bench为分析当前MLLMs的优缺点提供了关键见解,以推动更具有社交智能的AI的发展。数据集和代码可在https://kfq20.github.io/sivbench/获取。
SIV-Bench is a video benchmark dataset designed to rigorously evaluate the capabilities of Multimodal Large Language Models (MLLMs) in three core social intelligence tasks: Social Scene Understanding (SSU), Social State Reasoning (SSR), and Social Dynamic Prediction (SDP). This dataset contains 2,792 video clips and 8,792 meticulously generated question-answer pairs, which are derived from a pipeline where humans collaborate with LLMs. It was initially collected from TikTok and YouTube, covering a wide spectrum of video genres, presentation styles, and linguistic and cultural backgrounds. SIV-Bench also specifically offers an experimental setup for analyzing the impact of different textual cues, including raw on-screen text, added dialogues, or no text at all. Comprehensive experiments conducted on state-of-the-art MLLMs have demonstrated that these models perform excellently in handling SSU, but encounter substantial difficulties in SSR and SDP, with Relation Inference (RI) acting as a critical bottleneck. SIV-Bench provides critical insights for analyzing the strengths and limitations of current MLLMs, with the goal of promoting the development of AI endowed with stronger social intelligence. The dataset and code are accessible at https://kfq20.github.io/sivbench/.




