NJU-LINK/OmniCap-IF-54K
收藏资源简介:
OmniCap-IF-54K 是一个大规模指令调优数据集,旨在提升全模态视频字幕生成中的指令跟随能力。它包含54K个精心策划的视频-指令-响应三元组,涵盖格式约束、时间基础、视觉和音频内容约束以及视听协同。该数据集通过三阶段流程构建:视频筛选、约束感知指令合成和解耦响应生成,旨在训练模型生成有用的全视频字幕,同时遵守复杂的用户指定要求,如JSON模式、Markdown表格、时间戳格式、事件定位和跨模态推理。
OmniCap-IF-54K is a large-scale instruction-tuning dataset for improving instruction-following abilities in omni-modal video captioning. It contains 54K curated video-instruction-response triplets covering format constraints, temporal grounding, visual and audio content constraints, and audio-visual synergy. The dataset is constructed through a three-stage pipeline: video curation, constraint-aware instruction synthesis, and decoupled response generation. The resulting samples are designed to train models to produce useful omni-video captions while obeying complex user-specified requirements such as JSON schemas, Markdown tables, timestamp formats, event localization, and cross-modal reasoning.




