Dense Instruction Dataset
收藏资源简介:
Dense Instruction Dataset是由香港中文大学MMLab和英伟达等机构创建的一个密集指令数据集,旨在支持流式视频交互模型的训练。该数据集包含51,000条指令-答案对,每对都带有时间戳,模拟了流式视频交互的动态变化。数据集的创建过程结合了现有的密集字幕数据集,并通过启发式方法为每个词分配时间戳,确保模型在训练时能够模拟真实的流式交互场景。该数据集主要应用于流式视频交互领域,旨在提升大模态模型在动态视频环境中的交互能力和响应准确性。
Dense Instruction Dataset is a dense instruction dataset created by MMLab of The Chinese University of Hong Kong, NVIDIA and other institutions, aiming to support the training of streaming video interaction models. This dataset contains 51,000 instruction-answer pairs, each accompanied by timestamps, simulating the dynamic changes in streaming video interactions. The dataset was developed by combining existing dense caption datasets, and timestamps were assigned to each word via heuristic methods, ensuring that the model can simulate real streaming interaction scenarios during training. This dataset is primarily applied in the field of streaming video interaction, with the goal of enhancing the interactive ability and response accuracy of large multimodal models in dynamic video environments.




