SimLife
收藏资源简介:
SimLife长视频数据集是一个基于《模拟人生4》游戏模拟构建的长视频理解基准测试集,旨在评估模型在长时域视频和音频理解任务中的能力,如跨日记忆、时序推理、行为预测和反事实推理。数据集包含466个模拟日视频单元(每个约24分钟),这些单元被组织成106个视频链(其中77个包含用于音频基础推理的omni对话覆盖),以及381个评估任务文件,共计1439个多项选择题。每个视频单元配备结构化的事件日志(包含动作和对话事件)、密集的自然语言字幕(60秒窗口)和对话转录。任务设计为多日视频链场景,要求模型基于观察到的行为模式回答关于过去或未来行为的问题,问题类型包括直接预测、反事实推理和噪声反事实推理。当前版本完整发布了所有文本和结构化标注(包括任务、问题、答案、事件日志、字幕和对话转录),而视频和音频文件(约640GB)暂缓发布,但提供了两个完整单元(ID 1457和1417)作为示例。数据集适用于视觉问答、视频文本到文本和问答等任务,特别关注长视频、长上下文、时序推理和反事实推理等研究领域。
The SimLife Long Video Dataset is a long-form video understanding benchmark constructed through simulation of the game *The Sims 4*, designed to evaluate models' capabilities in long-temporal video and audio understanding tasks such as cross-day memory, temporal reasoning, behavior prediction, and counterfactual reasoning. The dataset contains 466 simulated daily video clips (each approximately 24 minutes long), which are organized into 106 video chains (77 of which include omni dialogue coverage for basic audio reasoning), along with 381 evaluation task files totaling 1,439 multiple-choice questions. Each video clip is equipped with structured event logs (covering action and dialogue events), dense natural language subtitles with a 60-second window, and full dialogue transcripts. The tasks are designed in multi-day video chain scenarios, requiring models to answer questions about past or future behaviors based on observed behavioral patterns, with question types including direct prediction, counterfactual reasoning, and noisy counterfactual reasoning. The current version fully releases all text and structured annotations including tasks, questions, answers, event logs, subtitles, and dialogue transcripts, while the video and audio files (approximately 640GB in total) are temporarily withheld from release, but two complete clips with IDs 1457 and 1417 are provided as examples. This dataset is applicable to tasks such as visual question answering (VQA), video-to-text generation, and question answering (QA), with a particular focus on research areas including long-form video, long context, temporal reasoning, and counterfactual reasoning.




