遇见数据集

trumancai/lco-compositional-train-gpt55-filtered

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

LCO Compositional Train — gpt-5.5 (vision-judge filtered) 是一个经过过滤的视觉语言数据集,基于trumancai/lco-compositional-train-gpt55数据集构建。原始数据集包含10002条记录,每条记录通过gpt-5视觉输入进行独立判断,依据三个严格标准:正样本标题正确描述图像、负样本标题不再匹配图像、以及变化发生在预期的组合轴上(如关系中的主语/宾语交换、属性绑定或顺序重排)。仅当所有标准都为真时,记录才被保留,最终保留8336条记录,保留率为83.3%。数据集涵盖三个类别:关系(2924条)、属性(2572条)和顺序(2840条),并包含硬负样本(6848条)和硬正样本(1488条)。数据集中图像来自MS-COCO 2014(CC-BY 4.0许可),标题来自COCO captions(CC-BY 4.0许可),硬负样本/正样本由OpenAI gpt-5.5生成,并由gpt-5判断。该数据集适用于图像文本到文本任务,支持组合性、对比学习和视觉语言研究。

LCO Compositional Train — gpt-5.5 (vision-judge filtered) is a filtered vision-language dataset derived from the trumancai/lco-compositional-train-gpt55 dataset. The original dataset consists of 10002 records, each independently judged by gpt-5 with vision input against three strict criteria: positive caption correctly describes the image, hard negative caption no longer matches the image, and the change is on the expected compositional axis (e.g., subject/object swap for relation, attribute binding for attribution, constituent reorder for order). A record is kept only if all three criteria are true, resulting in 8336 kept records (83.3% retention). The dataset includes three categories: relation (2924 records), attribution (2572 records), and order (2840 records), with hard negatives (6848 records) and hard positives (1488 records). Images are sourced from MS-COCO 2014 (CC-BY 4.0 license), captions from COCO captions (CC-BY 4.0 license), and hard negatives/positives are generated by OpenAI gpt-5.5 and judged by gpt-5. It is designed for image-text-to-text tasks, focusing on compositionality, contrastive learning, and vision-language research.

提供机构:
trumancai
二维码
社区交流群
二维码
科研交流群
商业服务