RUC-AIBOX/ClawGym-Bench
收藏资源简介:
ClawGym-Bench 是一个包含200个实例的诊断基准,专为Claw-style智能体设计。每个任务包括用户指令、模拟工作空间资源和一个任务特定的验证器。其中,156个任务使用基于代码的验证,44个任务采用混合验证(结合代码检查和基于规则的判断),混合验证的评分权重为0.7(代码验证)和0.3(规则验证)。该基准通过难度感知过滤和人工-LLM审查选择,涵盖六个工作空间基础类别:产品与协作、系统与自动化、分析与推理、内容与领域、规划与知识、软件开发。
ClawGym-Bench is a diagnostic benchmark comprising 200 instances, specifically tailored for Claw-style AI Agents. Each task consists of user instructions, simulated workspace resources, and a task-specific validator. Of these, 156 tasks utilize code-based validation, while 44 tasks adopt hybrid validation (combining code inspection and rule-based judgment), with the scoring weights of hybrid validation set as 0.7 for code validation and 0.3 for rule-based validation. This benchmark is curated via difficulty-aware filtering and human-LLM review, covering six fundamental workspace categories: Products and Collaboration, Systems and Automation, Analysis and Reasoning, Content and Domain, Planning and Knowledge, and Software Development.




