遇见数据集

ITBench-Trajectories

收藏
魔搭社区2026-07-14 更新2026-07-15 收录
官方服务:

资源简介:

# ITBench Trajectories This dataset contains complete execution trajectories of LLM agents using the [ITBench-SRE-Agent](https://github.com/itbench-hub/ITBench-SRE-Agent). It captures real agent reasoning, tool usage, and performance across multiple state-of-the-art language models tackling Site Reliability Engineering (SRE), Security & Compliance (CISO), and Financial Operations (FinOps) scenarios from the [ITBench benchmark](https://github.com/IBM/ITBench). ## Dataset Description ITBench Trajectories is a comprehensive collection of agent execution traces for IT automation tasks. Each trajectory includes the agent's investigation workflow, generated code, final output, and quantitative evaluation metrics. ### Dataset Summary - **105 complete agent trajectories** across 35 ITBench SRE scenarios (3 runs per scenario) - **1 state-of-the-art open source LLM**: OpenAI GPT-OSS-120B - **2 additional models (to be released)**: Google Gemini-3-Flash-Preview, MoonshotAI Kimi-K2-Thinking (each will include 3 runs per scenario) - **Detailed evaluation metrics** including precision, recall, F1 scores for root cause identification - **Rich context**: alerts, events, metrics, traces, logs, and Kubernetes object states ### What Are ITBench Scenarios? The scenarios in this dataset come from [ITBench](https://github.com/IBM/ITBench), an open-source benchmarking platform for evaluating AI agents on realistic IT automation tasks spanning multiple domains: **SRE (Site Reliability Engineering)**: Environment snapshots capturing observability data (logs, traces, metrics, alerts, events) from orchestrated Kubernetes environments where faults were injected. Task: Identify the faulty entity or resource based on the provided observability data and explain all firing alerts. See [ITBench SRE scenarios](https://github.com/IBM/ITBench/tree/main/scenarios/sre) for details. **CISO (Security & Compliance)**: Scenarios providing regulatory rules and natural-language Kubernetes/RHEL configurations. Task: Read the requirement and configuration, then generate the correct Kyverno policy (Kubernetes) or OPA policy (RHEL) to enforce compliance (to be added in future releases). **FinOps (Financial Operations)**: Synthetic data mimicking real-world cost and efficiency patterns. Task: Answer queries about anomalous cost changes by identifying the entity or resource responsible and explaining what caused the cost anomaly (to be added in future releases). ### Supported Tasks - **Agent Reasoning Analysis**: Study how LLM agents approach complex diagnostic and operational tasks - **Root Cause Analysis Evaluation**: Benchmark fault localization and propagation chain identification (SRE) - **Security & Compliance Assessment**: Evaluate compliance posture and security enforcement (CISO, future) - **Cost Optimization Analysis**: Analyze cloud spending and identify cost anomalies (FinOps, future) - **Code Generation**: Analyze Python code generated by agents for data analysis - **Multi-step Reasoning**: Examine agent decision-making across investigation phases - **Tool Use Patterns**: Understand how agents leverage observability data and automation tools ## Dataset Structure ``` ReAct-Agent-Trajectories/ ├── OpenAI-GPT-OSS-120B/ │ └── sre/ │ ├── Scenario-1/ │ │ ├── 1/ │ │ │ ├── agent_output.json # Agent's final diagnosis │ │ │ ├── session.jsonl # Complete execution trace │ │ │ ├── judge_output.json # Evaluation metrics │ │ │ └── code_generated_by_agent/ # Python scripts (optional) │ │ ├── 2/ │ │ └── 3/ │ ├── Scenario-2/ │ ├── Scenario-4/ │ └── ... (35 scenarios total) ├── Google-Gemini-3-Flash-Preview/ # To be released │ ├── sre/ │ ├── ciso/ # Future addition │ └── finops/ # Future addition └── MoonshotAI-Kimi-K2-Thinking/ # To be released ├── sre/ ├── ciso/ # Future addition └── finops/ # Future addition ``` ### Domains and Scenarios **SRE (Site Reliability Engineering)** The dataset currently includes **35 distinct ITBench SRE scenarios**. Currently, **3 runs per scenario** are available for the OpenAI GPT-OSS-120B model. Additional runs and models will be released in the future. **CISO (Security & Compliance)** CISO trajectories will be added in a future release. **FinOps (Financial Operations)** FinOps trajectories will be added in a future release. ## Data Collection ### Agent System Trajectories were collected using the [ITBench-SRE-Agent](https://github.com/itbench-hub/ITBench-SRE-Agent), a ReAct-style agent with: - **Tool Use**: Bash commands, Python code execution, file operations - **Observability Access**: Alerts, events, metrics, traces, logs - **Structured Output**: JSON diagnosis with entities, propagations, and explanations - **Multi-step Reasoning**: Phase-based investigation workflow ### Models Evaluated 1. **OpenAI GPT-OSS-120B** - OpenAI's 120B parameter open-source model (available now) 2. **Google Gemini-3-Flash-Preview** - Google's Gemini 3 Flash Preview (to be released) 3. **MoonshotAI Kimi-K2-Thinking** - Moonshot AI's Kimi K2 with extended reasoning (to be released) ### Evaluation Each diagnosis was automatically evaluated against ground truth using: - Entity identification (precision, recall, F1) - Propagation chain accuracy - Alert explanation completeness - Fault localization metrics (service and component level) - Ranked metrics for entity prioritization ## Use Cases - **Agent Reasoning Research**: Study how LLM agents approach complex diagnostic and operational tasks - **Benchmarking**: Compare fault localization, compliance assessment, and cost optimization across models - **Training Data**: Fine-tune models or train new agents for IT automation tasks - **Tool Use Analysis**: Understand how agents leverage observability data and automation tools - **Multi-step Reasoning**: Examine agent decision-making across investigation phases ## Citation If you use this dataset in your research, please cite: ```bibtex @misc{itbench-trajectories-2025, title={ITBench Agent Trajectories: LLM Agent Executions for IT Automation Tasks}, author={ITBench Team}, year={2025}, publisher={Hugging Face}, howpublished={\url{https://huggingface.co/datasets/itbench/itbench-trajectories}} } ``` ## Agent Implementation The agent that generated these trajectories is open source and available at: **[https://github.com/itbench-hub/ITBench-SRE-Agent](https://github.com/itbench-hub/ITBench-SRE-Agent)** The agent uses a ReAct (Reasoning + Acting) framework to iteratively investigate incidents by: 1. Reasoning about the incident using available observability data 2. Acting by executing tools (Bash, Python, file operations) 3. Refining hypotheses based on evidence 4. Constructing causal chains and explanations ## Additional Resources - **ITBench Benchmark**: [https://github.com/IBM/ITBench](https://github.com/IBM/ITBench) - **ITBench Scenarios**: - [SRE](https://github.com/IBM/ITBench/tree/main/scenarios/sre) - [CISO](https://github.com/IBM/ITBench/tree/main/scenarios/ciso) - [FinOps](https://github.com/IBM/ITBench/tree/main/scenarios/finops) - **ITBench-SRE-Agent**: [https://github.com/itbench-hub/ITBench-SRE-Agent](https://github.com/itbench-hub/ITBench-SRE-Agent)

提供机构:
maas
创建时间:
2026-01-18
二维码
社区交流群
二维码
科研交流群
商业服务