lmarena-ai/Arena-T2I-Hard
收藏资源简介:
Arena-T2I-Hard是一个包含310个提示的压力测试基准,用于评估文本到图像模型的忠实度(即提示跟随能力)。数据来源于真实且困难的用户请求,这些请求通常包含长文本、多实体、属性、空间关系、计数和风格约束。每个提示都被预先分解为依赖感知的有向无环图(DAG)的是/否问题,当评估图像时,如果父问题失败,则其子问题的得分将被清零。该基准在DPG-Bench和DSG等基准饱和时仍能保持区分性。数据集包括约13.9k个问题,其中约9.6k个为忠实度问题,约4.4k个为美学问题。
Arena-T2I-Hard is a 310-prompt stress benchmark for evaluating faithfulness (prompt-following) of text-to-image models, drawn from real, hard arena user requests — long, multi-entity prompts with attributes, spatial relations, counts, and stylistic constraints. Each prompt ships pre-decomposed into a dependency-aware DAG of yes/no questions; when scoring an image, failing a parent question zeroes out its descendants. The benchmark stays discriminative where DPG-Bench and DSG saturate. Across the 310 prompts there are ~13.9k questions (~9.6k faithfulness, ~4.4k aesthetics).




