AgentWorldBench
收藏资源简介:
# AgentWorldBench AgentWorldBench is a comprehensive evaluation benchmark for language world models, constructed from real-world observations of frontier model trajectories on established benchmarks such as Tool Decathlon, Terminal-Bench 1.0 & 2.0, and OSWorld-Verified. Every evaluation sample is paired with a ground-truth observation obtained from real environment execution, enabling reference-grounded scoring. AgentWorldBench evaluates world modeling quality by scoring each predicted environment observation on five dimensions — **Format**, **Factuality**, **Consistency**, **Realism**, and **Quality** — probing the reasoning, knowledge, and long-context capabilities required for faithful environment simulation. For more details, please refer to the [technical report](http://arxiv.org/abs/2606.24597) and the [blog post](https://qwen.ai/blog?id=qwen-agentworld). ## Benchmark Statistics | Domain | Samples | Avg. Turns | Description | |--------|--------:|-----------:|-------------| | MCP | 286 | 23.1 | API server responses: tool call results, database state, service protocols | | Search | 458 | 15.5 | Search engine results: URLs, snippets, rankings, page content | | Terminal | 354 | 26.7 | Command-line environment: shell output, file system state, process behavior | | SWE | 472 | 28.1 | IDE / code editing environment: git diff, test results, compilation errors | | Android | 200 | 37.8 | Android UI hierarchy changes after touch/gesture actions | | Web | 200 | 14.2 | Browser DOM state changes after user interactions | | OS | 200 | 12.7 | Desktop OS state: file system, window management, application behavior | | **Total** | **2,170** | **22.8** | | ## Data Format Each file is a per-domain JSONL (`{domain}_test.jsonl`). Each record is a single evaluation turn from a multi-turn environment trajectory. `prompt` and `response` are **parallel lists of length `turn_idx`**, representing the full conversation history up to and including the evaluated turn. The ground-truth observation for the current turn is always the **last element** `response[-1]`, while earlier elements provide context from preceding turns. ```json { "task": "terminal", "id": 267463494664789, "prompt": [ "### Turn 1\n**Action:**\n```json\n[{\"keystrokes\": \"ls -la\\n\"}]\n```", "### Turn 2\n**Action:**\n```json\n[{\"keystrokes\": \"cat README.md\\n\"}]\n```", "### Turn 3\n**Action:**\n```json\n[{\"keystrokes\": \"mkdir output\\n\"}]\n```" ], "response": [ "**Environment Observation:**\nroot@2b1e6f43cde5:/app# ls -la\ntotal 20\n...", "**Environment Observation:**\nroot@2b1e6f43cde5:/app# cat README.md\n...", "**Environment Observation:**\nroot@2b1e6f43cde5:/app# mkdir output\nroot@2b1e6f43cde5:/app#" ], "current_prompt": "### Turn 3\n**Action:**\n```json\n[{\"keystrokes\": \"mkdir output\\n\"}]\n```", "system_str": "# Role and Objective\n\nYou are a **Terminal World Model** ...", "turn_idx": 3, "total_turns": 151 } ``` **Fields:** | Field | Description | |-------|-------------| | `task` | Domain identifier (`mcp`, `search`, `terminal`, `swe`, `android`, `web`, `os`) | | `id` | Trajectory identifier (shared by all samples from the same trajectory) | | `prompt` | List of action prompts from turn 1 through `turn_idx`. `prompt[i]` is the action at turn `i+1` | | `response` | List of ground-truth observations from turn 1 through `turn_idx`. **`response[-1]` is the ground truth for the evaluated turn**; earlier elements are context | | `current_prompt` | The action prompt for the evaluated turn (same as `prompt[-1]`) | | `system_str` | The world model system prompt for this sample | | `turn_idx` | 1-indexed position of the evaluated turn | | `total_turns` | Total number of turns in the source trajectory | > **Note:** Each trajectory may appear as multiple records with different `turn_idx` values, each evaluating a different point in the trajectory. Container/session IDs (e.g., `root@2b1e6f43cde5`) are consistent within a trajectory but differ across trajectories, as each runs in its own environment. ## Evaluation We provide a standalone evaluation script in the [GitHub repository](https://github.com/QwenLM/Qwen-AgentWorld/tree/main/eval). The evaluation follows a three-step pipeline: ```bash cd eval # Step 1: Run world model inference python eval.py infer \ --data-dir ../AgentWorldBench \ --model-base-url http://localhost:8000/v1 \ --model-name Qwen/Qwen-AgentWorld-35B-A3B \ --output-dir ./results # Step 2: Run LLM judge scoring export OPENAI_API_KEY="your-api-key" python eval.py judge \ --predictions ./results/predictions.jsonl \ --judge-base-url https://api.openai.com/v1 \ --judge-model gpt-5.2-2025-12-11 \ --output-dir ./results # Step 3: Aggregate and display scores python eval.py score --predictions ./results/judged.jsonl ``` See the [GitHub README](https://github.com/QwenLM/Qwen-AgentWorld#evaluate-on-agentworldbench) for full setup instructions, deployment guides, and domain-specific system prompt templates. ## Citation ```bibtex @article{zuo2026qwen, title={Qwen-agentworld: language world models for general agents}, author={Zuo, Yuxin and Xiao, Zikai and Sheng, Li and Huang, Fei and Tu, Jianhong and Liu, Yuxuan and Tang, Tianyi and Hu, Xiaomeng and Su, Yang and Lan, Qingfeng and others}, journal={arXiv preprint arXiv:2606.24597}, year={2026} } ```



