BizGenEval
收藏资源简介:
BizGenEval 是一个用于评估图像生成模型在现实商业设计任务上表现的基准数据集。该数据集覆盖了5种文档类型(幻灯片、网页、海报、图表、科学图表)和4种能力维度(文本渲染、布局控制、属性绑定、知识推理),共包含20个评估任务。数据集提供了400个精心设计的提示词和8,000个检查清单问题(4,000个简单问题 + 4,000个困难问题)。每个数据样本包含以下字段:唯一标识符(id)、文档类型(domain)、能力维度(dimension)、目标宽高比(aspect_ratio)、参考分辨率(reference_image_wh)、生成提示词(prompt)、20个是/否检查问题(questions)、评估标签(eval_tag)以及简单和困难问题的索引(easy_qidxs, hard_qidxs)。该数据集适用于商业视觉内容生成任务的评估和模型性能测试。
BizGenEval is a benchmark dataset for evaluating the performance of image generation models on real-world commercial design tasks. This dataset covers 5 document types (slides, web pages, posters, charts, scientific charts) and 4 capability dimensions (text rendering, layout control, attribute binding, knowledge reasoning), with a total of 20 evaluation tasks. It provides 400 meticulously designed prompts and 8,000 checklist questions (4,000 easy questions and 4,000 difficult questions). Each data sample contains the following fields: unique identifier (id), document type (domain), capability dimension (dimension), target aspect ratio (aspect_ratio), reference resolution (reference_image_wh), generation prompt (prompt), 20 yes/no check questions (questions), evaluation tag (eval_tag), as well as the indexes of easy and difficult questions (easy_qidxs, hard_qidxs). This dataset is suitable for the evaluation of commercial visual content generation tasks and model performance testing.




