Neomi26/cc12m-images-011
收藏资源简介:
cc12m-images-011数据集是从CC12M数据集中提取的一个子集,包含90,000张图像,以单独的JPEG文件形式重新托管,便于通过浏览器直接访问。该数据集主要用于图像到文本任务,如图像标注或生成。图像URL遵循特定模式:https://huggingface.co/datasets/Neomi26/cc12m-images-011/resolve/main/{folder}/{key}.jpg。清单文件(manifest.parquet)包含多个列,如key、image_url、source_url、caption、width、height、original_width、original_height、shard、repo和path,提供了图像的元数据信息。该数据集属于Neomi26/cc12m-images-NNN系列的一部分,该系列共约131个仓库,并有一个全局索引数据集(Neomi26/cc12m-images-index)用于整体管理。原始数据集来源于pixparse/cc12m-wds。
The cc12m-images-011 dataset is a subset extracted from the CC12M dataset, containing 90,000 images that are re-hosted as individual JPEG files for direct browser-based access. This dataset is primarily designed for image-to-text tasks, such as image captioning or image generation. The image URLs follow a specific pattern: https://huggingface.co/datasets/Neomi26/cc12m-images-011/resolve/main/{folder}/{key}.jpg. The manifest file (manifest.parquet) includes multiple columns such as key, image_url, source_url, caption, width, height, original_width, original_height, shard, repo, and path, providing metadata information for the images. This dataset is part of the Neomi26/cc12m-images-NNN series, which consists of approximately 131 repositories, with a global index dataset (Neomi26/cc12m-images-index) used for overall management. The original dataset is sourced from pixparse/cc12m-wds.




