Awiny/Howto-Interlink7M
收藏资源简介:
--- license: apache-2.0 --- # Howto-Interlink7M ## 📙 Overview Howto-Interlink7M presents a unique interleaved video-text dataset, carefully derived from the raw video content of [Howto100M](https://www.di.ens.fr/willow/research/howto100m/). <img src="howto_interlink7m_ppl.png" width="75%" height="75%"> In the creation of this dataset, we turn **a long video into a vision-text interleaved documents** by BLIP2 (Img Captioner), GRIT (Img Detector), Whisper (ASR). Similar to [VLog](https://github.com/showlab/VLog). Then, we employed the **GPT-4** for an extensive **7 million** high-quality pretraining data. During this process, we meticulously filtered out clips containing sensitive or low-quality content. <img src="https://cdn-uploads.huggingface.co/production/uploads/64440be5af034cdfd69ca3a7/tCl0r7zasZwwV1qJF1OJN.png" width="50%" height="50%"> ## 📊 Statistics The statictis are listed below: | Split | Samples | Average Clips | Average Clip Length | Average Document Tokens | |---|---|---|---| --- | | Howto-Interlink7M_subset_w_all_clips_train.tsv | 276711 | 8.4 | 49.8 | 460.3 | | Howto-Interlink7M_subset_w_all_clips_val.tsv | 30746 | 8.4 | 49.8 | 460.2 | | Howto-Interlink7M_subset_w_sampled_clips_train.tsv | 660827 | 5.8 | 47.2 |319.4 | | Howto-Interlink7M_sbset_w_sampled_clips_val.tsv| 73426| 5.8 | 47.2 | 319.8 | |All| 1041710| 6.6 | 48.0 | 361.0| ## 🎨 Visualization  Please see [Youtube](https://www.youtube.com/watch?v=z3uOI6oInto) for more examples. ## 🏋️ Training Please refer to code [cosmo](https://github.com/showlab/cosmo/) for training details. ## Download Source Video ### 1. Download the README and All-in-One zip file: On the official website [HowTo100M](https://www.di.ens.fr/willow/research/howto100m/), locate the download links for the README and the All-in-One zip file. Extract the contents of the All-in-One zip file: ### 2. Inside the extracted folder, you should find the HowTo100M_v1.csv file. ### 3. In the CSV file, you will find a column named "video_id" which contains unique identifiers for each video. You can use youtube-dl or similar tools to download the videos using the video IDs listed in the CSV file. ## 🎓 Citation ``` @article{wang2024cosmo, title={COSMO: Contrastive Streamlined Multimodal Model with Interleaved Pre-Training}, author={Wang, Alex Jinpeng and Li, Linjie and Lin, Kevin Qinghong and Wang Jianfeng and Lin, Kevin and Yang, Zhengyuan and Wang, Lijuan and Shou, Mike Zheng}, journal={arXiv preprint arXiv:2401.00849}, year={2024} } ```
--- 许可证:Apache-2.0 --- # Howto-Interlink7M ## 📙 概述 Howto-Interlink7M 是一款独特的交错式视频-文本数据集,其全部素材均源自 [Howto100M](https://www.di.ens.fr/willow/research/howto100m/) 的原始视频内容。 <img src="howto_interlink7m_ppl.png" width="75%" height="75%"> 在本数据集的构建流程中,我们借助**BLIP2(图像字幕模型)**、**GRIT(图像检测模型)**与**Whisper(自动语音识别,Automatic Speech Recognition)**,将长视频转换为视觉-文本交错文档,该构建思路与 [VLog](https://github.com/showlab/VLog) 一脉相承。 随后,我们使用**GPT-4**生成了总计700万条高质量预训练数据。 在此过程中,我们会严格过滤包含敏感内容或低质量素材的视频片段。 <img src="https://cdn-uploads.huggingface.co/production/uploads/64440be5af034cdfd69ca3a7/tCl0r7zasZwwV1qJF1OJN.png" width="50%" height="50%"> ## 📊 统计信息 数据集的统计信息如下: | 数据集划分 | 样本量 | 平均片段数 | 平均片段时长 | 平均文档Token数 | |---|---|---|---|---| | Howto-Interlink7M_subset_w_all_clips_train.tsv | 276711 | 8.4 | 49.8 | 460.3 | | Howto-Interlink7M_subset_w_all_clips_val.tsv | 30746 | 8.4 | 49.8 | 460.2 | | Howto-Interlink7M_subset_w_sampled_clips_train.tsv | 660827 | 5.8 | 47.2 | 319.4 | | Howto-Interlink7M_subset_w_sampled_clips_val.tsv | 73426 | 5.8 | 47.2 | 319.8 | | 总计 | 1041710 | 6.6 | 48.0 | 361.0 | ## 🎨 可视化  更多示例可参阅 [Youtube](https://www.youtube.com/watch?v=z3uOI6oInto)。 ## 🏋️ 训练 训练细节请参阅代码仓库 [cosmo](https://github.com/showlab/cosmo/)。 ## 📥 原始视频下载 ### 1. 下载README文档与全量压缩包 前往官方网站 [HowTo100M](https://www.di.ens.fr/willow/research/howto100m/),获取README文档与全量压缩包的下载链接,解压该全量压缩包。 ### 2. 在解压后的文件夹中,你将找到 HowTo100M_v1.csv 文件。 ### 3. 在该CSV文件中,存在一个名为`video_id`的列,其中包含每个视频的唯一标识符。你可借助`youtube-dl`或同类工具,通过CSV文件中列出的视频ID下载对应视频。 ## 🎓 引用格式 @article{wang2024cosmo, title={COSMO: Contrastive Streamlined Multimodal Model with Interleaved Pre-Training}, author={Wang, Alex Jinpeng and Li, Linjie and Lin, Kevin Qinghong and Wang Jianfeng and Lin, Kevin and Yang, Zhengyuan and Wang, Lijuan and Shou, Mike Zheng}, journal={arXiv preprint arXiv:2401.00849}, year={2024} }
Howto-Interlink7M
📙 概述
Howto-Interlink7M 是一个独特的视频-文本交错数据集,源自 Howto100M 的原始视频内容。该数据集通过 BLIP2(图像描述器)、GRIT(图像检测器)和 Whisper(自动语音识别)将长视频转化为视觉-文本交错文档。随后,使用 GPT-4 生成了 700 万 高质量预训练数据,过程中仔细过滤了包含敏感或低质量内容的片段。
📊 统计数据
以下是数据集的统计信息:
| 分割 | 样本数 | 平均片段数 | 平均片段长度 | 平均文档令牌数 |
|---|---|---|---|---|
| Howto-Interlink7M_subset_w_all_clips_train.tsv | 276711 | 8.4 | 49.8 | 460.3 |
| Howto-Interlink7M_subset_w_all_clips_val.tsv | 30746 | 8.4 | 49.8 | 460.2 |
| Howto-Interlink7M_subset_w_sampled_clips_train.tsv | 660827 | 5.8 | 47.2 | 319.4 |
| Howto-Interlink7M_sbset_w_sampled_clips_val.tsv | 73426 | 5.8 | 47.2 | 319.8 |
| 总计 | 1041710 | 6.6 | 48.0 | 361.0 |
🎓 引用
@article{wang2024cosmo, title={COSMO: Contrastive Streamlined Multimodal Model with Interleaved Pre-Training}, author={Wang, Alex Jinpeng and Li, Linjie and Lin, Kevin Qinghong and Wang Jianfeng and Lin, Kevin and Yang, Zhengyuan and Wang, Lijuan and Shou, Mike Zheng}, journal={arXiv preprint arXiv:2401.00849}, year={2024} }




