neulab/pace-bench
收藏资源简介:
PACE-Bench是由PACE(代理能力评估代理)产生的具体代理基准。它是PACE选择框架在4个代理目标(GAIA、SWE-Bench Verified、SWE-Bench Multimodal和SWT-Bench)上运行后输出的数据集,而非独立手动策划的基准。该数据集包含从12个非代理基准中自动选择的紧凑源实例集合,用于预测模型在代理目标上的性能。每个代理目标的代理由约100个源实例组成,每个实例带有其在PACE线性预测器中的回归权重。评估LLM代理在如SWE-Bench和GAIA等基准上通常成本高昂且缓慢,而PACE通过选择廉价非代理评估中的实例,其聚合分数能可靠预测代理性能。数据集文件包括每个代理目标的JSONL文件和一个图像文件夹,用于多模态实例。实例源自多个上游基准,内容遵循原始许可。
PACE-Bench is a concrete proxy benchmark generated by PACE (Proxy Capability Evaluation Agent). It is a dataset output by the PACE selection framework after running on four proxy goals: GAIA, SWE-Bench Verified, SWE-Bench Multimodal, and SWT-Bench, rather than an independently manually curated benchmark. This dataset comprises a compact set of source instances automatically selected from 12 non-proxy benchmarks, tailored for predicting model performance on proxy goals. For each proxy goal, the corresponding evaluation setup consists of approximately 100 source instances, each with its regression weight in the PACE linear predictor. Evaluating LLM agents on benchmarks such as SWE-Bench and GAIA is typically costly and time-consuming, whereas PACE selects instances from low-cost non-proxy evaluations, and its aggregated scores can reliably predict proxy performance. The dataset files include JSONL files for each proxy goal and an image folder for multimodal instances. All instances are sourced from multiple upstream benchmarks, and their content is subject to their original licenses.




