StreamingCoT
收藏资源简介:
StreamingCoT是一个针对流式视频问答和多模态思维链推理的大规模数据集。该数据集通过严格的分层流程构建,整合了时间分割、动态问答生成和多模态证据定位。数据集的构建包括多阶段验证,确保时空一致性和推理完整性。数据集通过YouTube官方API收集了10,288个短视频,并通过多模态过滤机制筛选出5,745个高质量视频。此外,数据集还采用了分层视频密集字幕框架,通过自适应时间分割和上下文感知叙述生成,解决了视频问答中答案的动态演变问题。
StreamingCoT is a large-scale dataset for streaming video question answering and multimodal chain-of-thought reasoning. It is built through a rigorous hierarchical construction pipeline that integrates temporal segmentation, dynamic question answering generation, and multimodal evidence localization. Multi-stage validation is incorporated during the dataset’s construction to ensure spatio-temporal consistency and reasoning completeness. A total of 10,288 short videos were collected via the official YouTube API, and 5,745 high-quality videos were screened out using a multimodal filtering mechanism. Furthermore, the dataset adopts a hierarchical dense video captioning framework, which addresses the dynamic evolution of answers in video question answering tasks through adaptive temporal segmentation and context-aware narrative generation.




