evaleval/auto-benchmarkcards
收藏资源简介:
该数据集包含用于AI评估基准的BenchmarkMetadataCards,这些卡片是自动生成的。BenchmarkCards是结构化的JSON文档,描述了基准测试的内容、工作原理及其局限性。它们涵盖了基准测试的目标、目标受众、数据源、方法、指标、局限性、伦理考量和相关AI风险等字段。数据集包含44张卡片,涵盖单个基准和复合基准套件,存储在benchmark-metadata.json文件和cards/文件夹中。卡片遵循IBM的AI Atlas Nexus的BenchmarkMetadataCard模式。复合卡片包含contains字段列出其子基准,而单个卡片包含appears_in字段链接到它们所属的任何父套件。卡片生成过程涉及从多个来源提取信息,并使用LLM将这些输入组合成结构化卡片。数据集目前是原型阶段,可能存在错误或不完整字段,建议进行人工审查。
This dataset contains automatically generated BenchmarkMetadataCards intended for AI evaluation benchmarks. BenchmarkMetadataCards (abbreviated as BenchmarkCards hereinafter) are structured JSON documents that elaborate on the content, operational principles, and limitations of the corresponding benchmarks. These cards include standardized fields such as the benchmark's objectives, target audience, data sources, methodologies, evaluation metrics, limitations, ethical considerations, and associated AI risks. The dataset consists of 44 such cards covering both standalone benchmarks and composite benchmark suites, and is stored in the benchmark-metadata.json file and the cards/ directory. All cards adhere to the BenchmarkMetadataCard schema defined by IBM's AI Atlas Nexus. Composite cards feature a "contains" field that lists their subordinate sub-benchmarks, while individual standalone cards include an "appears_in" field that links to any parent benchmark suites they belong to. The card generation pipeline involves extracting information from multiple sources, then synthesizing these collected inputs into structured BenchmarkMetadataCards using Large Language Models (LLMs). The dataset is currently in the prototype phase, which may contain errors or incomplete fields, thus manual review is highly recommended for practical applications.




