TAM Bench
收藏资源简介:
TAM Bench是一个为评估基于LLM的代理在端到端ML任务上的综合能力而设计的多样化、现实化和结构化的基准。该基准具有三个关键创新:(1)一个基于浏览器自动化和LLM的任务获取系统,自动从Kaggle、AIcrowd和Biendata等平台收集和结构化ML挑战,涵盖多种任务类型和数据模态;(2)一个基于排行榜的难度建模机制,使用参与者数量和分数分布来估计任务复杂性,实现可扩展和客观的任务校准;(3)一个多维度评估框架,结合性能、格式合规性、约束遵守和任务泛化。基于150个精心策划的AutoML任务,我们构建了三个不同大小的基准子集,包括Lite、Medium和Full,以适应不同的评估场景。
TAM Bench is a diverse, realistic, and structured benchmark designed to evaluate the comprehensive capabilities of LLM-based agents on end-to-end machine learning (ML) tasks. This benchmark features three core innovations: (1) A browser automation and LLM-powered task acquisition system that automatically collects and structures ML challenges from platforms such as Kaggle, AIcrowd, and Biendata, covering diverse task types and data modalities; (2) A leaderboard-based difficulty modeling mechanism that estimates task complexity using the number of participants and score distributions, enabling scalable and objective task calibration; (3) A multi-dimensional evaluation framework that integrates performance, format compliance, constraint adherence, and task generalization. Based on 150 carefully curated AutoML tasks, we constructed three benchmark subsets of varying sizes, namely Lite, Medium, and Full, to accommodate different evaluation scenarios.
- 1Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization复旦大学 · 2025年



