LAION-400M
收藏资源简介:
LAION-400M是由尤利希超级计算中心等机构创建的一个包含4亿对图像-文本数据的大型公开数据集。该数据集通过筛选自Common Crawl的图像及其对应的文本描述,并应用CLIP模型进行过滤,确保数据的质量和相关性。数据集创建过程中,采用了分布式处理和单节点后处理相结合的方法,以高效地从庞大的原始数据中提取和整理出高质量的图像-文本对。LAION-400M的应用领域广泛,主要用于训练多模态语言-视觉模型,如DALL-E和CLIP,以支持零样本或少量样本学习,解决图像和文本间的语义匹配问题。
LAION-400M is a large-scale public dataset containing 400 million image-text pairs, created by institutions including the Jülich Supercomputing Centre. This dataset is derived from images and their accompanying textual descriptions extracted from Common Crawl, and further filtered using the CLIP model to guarantee data quality and relevance. During the dataset construction, a hybrid approach combining distributed processing and single-node post-processing was adopted to efficiently extract and curate high-quality image-text pairs from the voluminous raw dataset. LAION-400M boasts a wide range of applications, primarily used for training multimodal language-vision models such as DALL-E and CLIP, to enable zero-shot or few-shot learning and tackle the semantic matching issue between images and text.

- 1LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs尤利希超级计算中心(JSC)研究中心尤利希(FZJ) · 2021年



