遇见数据集

TIGER-Lab/ClawBenchV2Trace

收藏
Hugging Face2026-05-25 更新2026-06-14 收录
官方服务:

资源简介:

ClawBench V2 Traces数据集是ClawBench基准测试V2版本中每个模型运行的完整执行轨迹。该数据集作为TIGER-Lab/ClawBench(任务定义)和NAIL-Group/ClawBenchV1Trace(V1轨迹)的配套数据集,发布了所有V2模型评估的原始执行数据。每个运行对应一个目录(按任务×模型×尝试组织),包含屏幕录制、网络捕获、浏览器操作、智能体推理和最终拦截的请求。这使得用户可以在不重新运行智能体的情况下,基于这些轨迹进行重新评分、调试或构建新的评估器。数据集布局包括每个运行目录中的文件,如recording.mp4(完整会话录制)、requests.jsonl(网络请求/响应)、actions.jsonl(浏览器操作流)、agent-messages.jsonl(智能体LLM推理轨迹)、interception.json(最终拦截的HTTP请求)、judge.json(LLM评估器对拦截负载的判决)等。数据集涵盖了多个模型(如claude-opus-4-7、glm-5.1、gpt-5.5、deepseek-v4-pro等),并提供了下载和评分指南,支持用户通过命令行工具下载特定模型或任务的轨迹,并使用两阶段评分标准(拦截和LLM评估)进行重新评估。数据集基于Apache 2.0许可证发布,并引用了相关研究论文。

ClawBench V2 Traces is a dataset containing full execution traces for every V2 model run scored on the ClawBench benchmark. It serves as a companion to TIGER-Lab/ClawBench (task definitions) and NAIL-Group/ClawBenchV1Trace (V1 traces). This dataset publishes the raw execution data for every V2 model run evaluated, with one directory per (task × model × attempt), each including screen recording, network capture, browser actions, agent reasoning, and the final intercepted request. This allows anyone to re-grade, debug, or build new evaluators on top of these traces without re-running the agent. The dataset layout includes files such as recording.mp4 (full session recording), requests.jsonl (network HTTP requests/responses), actions.jsonl (browser action stream), agent-messages.jsonl (agent LLM reasoning trace), interception.json (final intercepted HTTP request), judge.json (LLM judge verdict on the intercepted payload), and others. It covers multiple models (e.g., claude-opus-4-7, glm-5.1, gpt-5.5, deepseek-v4-pro, etc.) and provides download and scoring instructions, enabling users to download traces for specific models or tasks via command-line tools and re-evaluate using a two-stage rubric (interception and LLM judge). The dataset is licensed under Apache 2.0 and cites the associated research paper.

提供机构:
TIGER-Lab
二维码
社区交流群
二维码
科研交流群
商业服务