AGC-Bench
收藏资源简介:
AGC-Bench是由宾夕法尼亚州立大学等多家机构联合构建的人工通用创造力元基准数据集,旨在系统评估大语言模型在跨领域创造性任务中的表现。该数据集包含78个精选子集,涵盖头脑风暴、问题解决、STEM、叙事、比喻语言和幽默等六大文本领域,并包含多模态视觉与设计任务,数据源自对3,101篇文献的系统性综述与筛选。其创建过程遵循PRISMA标准,通过自动化预筛选与人工双重审核确保数据质量,并采用HELM兼容的标准化评估框架进行整合。该数据集主要应用于人工智能创造力研究,旨在探究大语言模型的创造性能力是否具有领域普适性,为解决人工通用创造力这一核心问题提供规模化评估基础设施。
AGC-Bench is a meta-benchmark dataset for artificial general creativity, jointly constructed by The Pennsylvania State University and multiple other institutions. It aims to systematically evaluate the performance of large language models (LLMs) on cross-domain creative tasks. This dataset includes 78 curated subsets covering six textual domains: brainstorming, problem-solving, STEM, narrative, figurative language, and humor, as well as multimodal visual and design tasks. The data is derived from a systematic review and screening of 3,101 academic papers. Its development follows the PRISMA guidelines, ensuring data quality through automated pre-screening and dual manual review, and integrates a standardized evaluation framework compatible with HELM. This dataset is primarily utilized in artificial intelligence creativity research, aiming to investigate whether the creative capabilities of LLMs exhibit domain generality, thus providing a scalable evaluation infrastructure for addressing the core issue of artificial general creativity.

- 1AGC-Bench: Measuring Artificial General Creativity宾夕法尼亚州立大学; 麻省大学洛厄尔分校; 阿姆斯特丹大学; 亚马逊AGI; 达特茅斯学院 · 2026年



