pixparse/cc12m-wds
收藏资源简介:
--- license: other license_name: conceptual-12m license_link: LICENSE task_categories: - image-to-text size_categories: - 10M<n<100M --- # Dataset Card for Conceptual Captions 12M (CC12M) ## Dataset Description - **Repository:** [Conceptual 12M repository](https://github.com/google-research-datasets/conceptual-12m) - **Paper:** [Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts](https://arxiv.org/abs/2102.08981) - **Point of Contact:** [Conceptual Captions e-mail](mailto:conceptual-captions@google.com) ### Dataset Summary Conceptual 12M (CC12M) is a dataset with 12 million image-text pairs specifically meant to be used for visionand-language pre-training. Its data collection pipeline is a relaxed version of the one used in Conceptual Captions 3M (CC3M). ### Usage This instance of Conceptual Captions is in [webdataset](https://github.com/webdataset/webdataset/commits/main) .tar format. It can be used with webdataset library or upcoming releases of Hugging Face `datasets`. ...More Detail TBD ### Data Splits This dataset was downloaded using img2dataset. Images resized on download if shortest edge > 512 to shortest edge = 512. #### Train * `cc12m-train-*.tar` * Downloaded on 2021/18/22 * 2176 shards, 10968539 samples ## Additional Information ### Dataset Curators Soravit Changpinyo, Piyush Sharma, Nan Ding and Radu Soricut. ### Licensing Information The dataset may be freely used for any purpose, although acknowledgement of Google LLC ("Google") as the data source would be appreciated. The dataset is provided "AS IS" without any warranty, express or implied. Google disclaims all liability for any damages, direct or indirect, resulting from the use of the dataset. ### Citation Information ```bibtex @inproceedings{changpinyo2021cc12m, title = {{Conceptual 12M}: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts}, author = {Changpinyo, Soravit and Sharma, Piyush and Ding, Nan and Soricut, Radu}, booktitle = {CVPR}, year = {2021}, } ```
license: 其他 license_name: conceptual-12m license_link: LICENSE task_categories: - 图像到文本 size_categories: - 1000万<样本量<1亿 # 概念性字幕1200万(CC12M)数据集卡片 ## 数据集说明 - **仓库地址**:[Conceptual 12M 仓库](https://github.com/google-research-datasets/conceptual-12m) - **相关论文**:[概念性字幕1200万:推进网页级图像-文本预训练以识别长尾视觉概念](https://arxiv.org/abs/2102.08981) - **联络方式**:[概念性字幕官方邮箱](mailto:conceptual-captions@google.com) ### 数据集概述 概念性字幕1200万(CC12M)是一个包含1200万图像-文本对的数据集,专为视觉-语言预训练任务设计。其数据收集流程是概念性字幕300万(CC3M)所使用流程的简化版本。 ### 使用方式 本版本的概念性字幕数据集采用[webdataset](https://github.com/webdataset/webdataset/commits/main)格式的.tar打包文件。可通过webdataset库或即将推出的拥抱脸(Hugging Face)`datasets`库进行使用。 ...更多细节待补充 ### 数据划分 本数据集通过img2dataset工具下载。下载过程中,若图像最短边大于512像素,则将其最短边调整为512像素。 #### 训练集 * `cc12m-train-*.tar` * 下载时间:2021/18/22 * 共2176个分片,包含10968539条样本 ## 补充信息 ### 数据集维护者 Soravit Changpinyo, Piyush Sharma, Nan Ding and Radu Soricut. ### 许可信息 本数据集可免费用于任何用途,若能注明谷歌有限责任公司(Google LLC,简称"谷歌")为数据来源,将不胜感激。本数据集按"现状"提供,不附带任何明示或暗示的担保。谷歌对因使用本数据集所导致的任何直接或间接损害不承担任何责任。 ### 引用信息 bibtex @inproceedings{changpinyo2021cc12m, title = {{Conceptual 12M}: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts}, author = {Changpinyo, Soravit and Sharma, Piyush and Ding, Nan and Soricut, Radu}, booktitle = {CVPR}, year = {2021}, }
数据集卡片 for Conceptual Captions 12M (CC12M)
数据集描述
- 数据集概述: Conceptual 12M (CC12M) 是一个包含1200万张图像-文本对的数据集,专门用于视觉和语言预训练。其数据收集流程是Conceptual Captions 3M (CC3M)的一个宽松版本。
使用方法
该版本的Conceptual Captions以webdataset .tar格式提供。可以使用webdataset库或即将发布的Hugging Face datasets进行使用。
数据分割
该数据集使用img2dataset下载,下载时如果最短边大于512,则调整为最短边为512。
训练集
cc12m-train-*.tar- 下载日期:2021/18/22
- 2176个分片,10968539个样本
附加信息
数据集策展人
Soravit Changpinyo, Piyush Sharma, Nan Ding 和 Radu Soricut。
许可信息
该数据集可自由用于任何目的,尽管对Google LLC ("Google")作为数据源的认可将受到赞赏。数据集以“AS IS”形式提供,没有任何明示或暗示的保证。Google不承担使用该数据集导致的任何直接或间接损害的责任。
引用信息
bibtex @inproceedings{changpinyo2021cc12m, title = {{Conceptual 12M}: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts}, author = {Changpinyo, Soravit and Sharma, Piyush and Ding, Nan and Soricut, Radu}, booktitle = {CVPR}, year = {2021}, }




