WorldGenBench
收藏资源简介:
WorldGenBench是一个用于评估文本到图像生成模型在理解世界知识和进行隐式推理能力方面的基准。数据集由1072个文本到图像的提示组成,分为人文和自然两个领域。每个提示都配有一个知识清单,用于评估模型生成图像的准确性。该数据集旨在解决现有模型在处理需要丰富世界知识和隐式推理的提示时的局限性问题。
WorldGenBench is a benchmark designed to evaluate the capability of text-to-image generation models in comprehending world knowledge and conducting implicit reasoning. The dataset comprises 1072 text-to-image prompts, categorized into two domains: humanities and natural sciences. Each prompt is accompanied by a knowledge checklist to assess the accuracy of the images generated by the models. This dataset aims to address the limitations of existing text-to-image models when handling prompts that require rich world knowledge and implicit reasoning.
WorldGenBench 数据集概述
基本信息
- 数据集名称: WorldGenBench
- 开发团队:
- Daoan Zhang1, Che Jiang3, Ruoshi Xu3, Biaoxiang Chen3, Zijian Jin4, Yutian Lu5
- Jianguo Zhang3, Liang Yong2, Jiebo Luo1, Shengda Luo2,3
- 机构: 1University of Rochester, 2Chinese Medicine Guangdong Laboratory, 3Southern University of Science and Technology
- 4New York University, 5Datawhale org.
数据集简介
- 目的: 评估文本到图像(T2I)生成模型的世界知识基础和隐式推理能力
- 特点:
- 覆盖人文和自然两大领域
- 提出"知识检查表分数"(Knowledge Checklist Score)作为结构化评估指标
- 实验发现:
- 扩散模型在开源方法中领先
- GPT-4o等专有自回归模型展现出更强的推理和知识整合能力
数据集内容
- 人文领域:
- 覆盖244个国家/地区
- 每个国家3个提示,共732个提示
- 主题: 历史、文化等
- 自然领域:
- 6个学科(天文学、物理学等)
- 共340个评估提示
- 质量控制: 所有提示经过人工验证以确保事实准确性和逻辑一致性
评估结果
- 人文领域表现最佳模型:
- GPT-4o (平均分24.46)
- HiDream-l1-Full (平均分16.68)
- SDv3.5-Large (平均分12.57)
- 自然领域表现最佳模型:
- GPT-4o (平均分19.61)
- Ideogram 2.0 (平均分9.34)
- SDv3.5-Large (平均分7.93)
数据示例
- 提示示例:
"In December 1982, deep in the rainforest of the Moumba province in western Gabon, a honey collector from the Miéné tribe was engaged in the traditional collection of wild honey..."




