foragi/try-v2
收藏资源简介:
该数据集是一个多模态视频问答数据集,设计用于评估模型在多种任务上的性能。它包含以下特征:id(标识符)、question_text(问题文本)、answer1和answer2(两个答案选项)、reminder1和reminder2(提醒信息)、video_type(视频类型)、video_duration(视频时长)、video(视频数据,未解码)和question_audio(问题音频)。数据集分为9个子集,包括PR_correction(校正任务)、PR_event_reminder(事件提醒任务)、PR_post_event_reminder(后事件提醒任务)、RTP_world_knowledge(世界知识任务)、RTP_counting(计数任务)、RTP_fine_grained_movement(细粒度运动任务)、RTP_interaction_relation(交互关系任务)、RTP_OCR(光学字符识别任务)和RTP_Omni(综合任务)。每个子集有5个示例,总数据集大小约为917MB,适用于视频理解和推理研究。
This dataset is a multimodal video question answering dataset designed to evaluate model performance across various tasks. It includes the following features: id (identifier), question_text (question text), answer1 and answer2 (two answer options), reminder1 and reminder2 (reminder messages), video_type (video type), video_duration (video duration), video (undecoded video data), and question_audio (question audio). The dataset is divided into 9 subsets, namely PR_correction (correction task), PR_event_reminder (event reminder task), PR_post_event_reminder (post-event reminder task), RTP_world_knowledge (world knowledge task), RTP_counting (counting task), RTP_fine_grained_movement (fine-grained movement task), RTP_interaction_relation (interaction relation task), RTP_OCR (Optical Character Recognition task), and RTP_Omni (comprehensive task). Each subset contains 5 examples, with a total dataset size of approximately 917 MB, and it is suitable for research on video understanding and reasoning.




