Dog100K
收藏资源简介:
Dog100K 是一个大规模、高质量的狗图像-文本对齐数据集,包含 103,508 个图像-文本对。该数据集专为图像-文本检索、多模态学习和条件图像生成等任务而设计。数据通过多阶段流程构建,包括从多样来源收集狗图像、进行分辨率、相关性和多样性方面的质量过滤、为每张图像标注包含品种、动作、场景以及是否有人类或多只狗存在的细粒度自然语言描述,并进行准确性验证。图像以 JPEG 格式存储,注释以 JSONL 格式存储,每个样本包含文件名(filename)、是否有人类(has_human)、是否有多个狗(multiple_dogs)、简要场景描述(scene)和详细的自然语言描述(description)五个字段。数据集覆盖了多样的犬种、姿态、场景、光照条件和背景,旨在增强模型的泛化能力。该数据集开源,可用于学术研究和工业应用,适用于图像-文本检索、图像描述生成、基于文本的条件图像生成(如 DiT、Stable Diffusion)以及多模态对比学习(如 CLIP、BLIP)等任务。
Dog100K is a large-scale, high-quality dog image-text alignment dataset consisting of 103,508 image-text pairs. This dataset is specifically developed for tasks such as image-text retrieval, multimodal learning, and conditional image generation. The dataset is built through a multi-stage pipeline, which involves collecting dog images from diverse sources, conducting quality filtering based on resolution, relevance and diversity, annotating each image with fine-grained natural language descriptions covering breed, action, scene, as well as the presence of humans or multiple dogs, and performing accuracy validation on the annotations. Images are stored in JPEG format, while the annotations are saved in JSONL format. Each sample includes five fields: filename, has_human, multiple_dogs, brief scene description (scene), and detailed natural language description (description). The dataset covers diverse dog breeds, poses, scenes, lighting conditions and backgrounds, with the goal of enhancing the generalization capability of models. This open-source dataset is accessible for both academic research and industrial applications, and is suitable for tasks including image-text retrieval, image captioning, text-conditioned image generation (e.g., DiT, Stable Diffusion), and multimodal contrastive learning (e.g., CLIP, BLIP).





