T2I-CoReBench
收藏资源简介:
T2I-CoReBench是一个用于评估文本到图像模型生成能力的基准数据集,包含组合和推理两个基本生成能力的12个维度。数据集平均提示长度为170个标记,平均问题数量为12.5个。该数据集旨在评估模型在复杂场景下的表现,并包含从简单到高复杂度的不同级别的任务。
T2I-CoReBench is a benchmark dataset for evaluating the generative capabilities of text-to-image models, encompassing 12 dimensions derived from its two fundamental generative abilities: composition and reasoning. The average prompt length of this dataset is 170 tokens, with an average of 12.5 questions per prompt. This dataset is designed to assess model performance in complex scenarios, and includes tasks ranging from simple to high-complexity levels.
T2I-CoReBench 数据集概述
基本信息
- 许可证: Apache-2.0
- 任务类别: 文本到图像
- 语言: 英语
- 标签: 文本到图像、评估、组合、推理
- 数据规模: 1K 到 10K 之间
- 官方名称: T2I-CoReBench
数据集配置
- 默认配置:
- 组合分割:
- C.MI: splits/C-MI.jsonl
- C.MA: splits/C-MA.jsonl
- C.MR: splits/C-MR.jsonl
- C.TR: splits/C-TR.jsonl
- 推理分割:
- R.LR: splits/R-LR.jsonl
- R.BR: splits/R-BR.jsonl
- R.HR: splits/R-HR.jsonl
- R.PR: splits/R-PR.jsonl
- R.GR: splits/R-GR.jsonl
- R.AR: splits/R-AR.jsonl
- R.CR: splits/R-CR.jsonl
- R.RR: splits/R-RR.jsonl
- 组合分割:
数据集特点
- 评估维度: 涵盖 12 个评估维度,分为组合和推理两个基本生成能力
- 组合能力:
- MI
- MA
- MR
- TR
- 推理能力:
- 演绎推理: LR、BR、HR、PR
- 归纳推理: GR、AR
- 溯因推理: CR、RR
- 复杂度: 平均提示长度为 170 个 token,平均包含 12.5 个问题
相关资源
- 项目页面: https://t2i-corebench.github.io/
- 论文: arXiv:2509.03516
- 数据集: https://huggingface.co/datasets/lioooox/T2I-CoReBench
- 图像数据: https://huggingface.co/datasets/lioooox/T2I-CoReBench-Images
- 代码: https://github.com/KwaiVGI/T2I-CoReBench
引用信息
bibtex @article{li2025easier, title={Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?}, author={Li, Ouxiang and Wang, Yuan and Hu, Xinting and Huang, Huijuan and Chen, Rui and Ou, Jiarong and Tao, Xin and Wan, Pengfei and Feng, Fuli}, journal={arXiv preprint arXiv:2509.03516}, year={2025} }




