Claw-Eval-Live
收藏资源简介:
Claw-Eval-Live是由香港中文大学等机构联合构建的动态工作流智能体评估基准,其核心特色在于通过ClawHub Top-500技能等公开需求信号动态生成任务。当前版本包含105项任务,涵盖商业服务流程和本地工作空间修复两大场景,涉及18个受控服务及沙箱环境。数据集通过信号聚类、权重分配、种子扩展等五阶段流程构建,采用混合式行动锚定评分机制,旨在解决传统静态基准与真实工作流需求脱节的问题,为LLM智能体在多系统协同、状态变更验证等复杂场景下的性能评估提供科学标准。
Claw-Eval-Live is a dynamic workflow AI Agent evaluation benchmark jointly developed by The Chinese University of Hong Kong and other institutions. Its core feature is dynamically generating tasks via public demand signals such as ClawHub Top-500 skills. The current version includes 105 tasks, covering two scenarios: commercial service processes and local workspace repair, involving 18 controlled services and sandbox environments. The dataset is constructed through a five-stage workflow encompassing signal clustering, weight assignment, seed expansion and other procedures. It adopts a hybrid action-anchored scoring mechanism, aiming to address the mismatch between traditional static benchmarks and real-world workflow requirements, and provide a scientific standard for performance evaluation of LLM-powered AI Agents in complex scenarios such as multi-system collaboration and state change verification.




