MineExplorer
收藏资源简介:
<div align="center"> <img src="https://raw.githubusercontent.com/Jometeorie/MineExplorer/main/figures/longcat-logo-full-20260504.png" height="68" alt="LongCat"/> <h1>MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft</h1> <p> Tianjie Ju · Yueqing Sun · Zheng Wu · Wei Zhang · Yaqi Huo · Xi Su · Qi Gu · Xunliang Cai · Gongshen Liu · Zhuosheng Zhang </p> <p> <a href="https://arxiv.org/abs/2605.30931"><img src="https://img.shields.io/badge/arXiv-2605.30931-b31b1b.svg" alt="arXiv"/></a> <a href="https://github.com/Jometeorie/MineExplorer"><img src="https://img.shields.io/badge/GitHub-MineExplorer-555555.svg?logo=github" alt="GitHub"/></a> <a href="https://huggingface.co/spaces/meituan-longcat/MineExplorer-Leaderboard"><img src="https://img.shields.io/badge/🤗-Leaderboard-blue.svg" alt="Leaderboard"/></a> <img src="https://img.shields.io/badge/License-MIT-555555.svg" alt="MIT"/> </p> </div> --- ## Abstract Multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and action generation. However, their ability to sustain exploration in dynamic open worlds remains unclear. Existing embodied and game-based benchmarks often compress interaction into short-horizon tasks or entangle success with domain-specific game mechanics. In this paper, we introduce MineExplorer benchmark for evaluating open-world exploration capabilities of MLLM agents in Minecraft. We first filter atomic tasks whose solutions rely heavily on Minecraft-specific knowledge to better reflect general open-world reasoning. Then we organize the benchmark around a ReAct-style capability formulation and compose atomic tasks into implicit multi-hop tasks. To further construct reliable instances, MineExplorer uses a multi-agent synthesis workflow that jointly designs task graphs, sandbox scenes, and rule-based milestone evaluators. Human evaluation shows that the multi-agent synthesis workflow produces significantly more reliable instances than a single-agent baseline. Experiments with advanced MLLM agents show that open-world exploration remains challenging, as strong models can handle many single-hop tasks but degrade sharply when hidden prerequisites must be coordinated over longer trajectories. Further analysis finds that task difficulty tracks agent completion, and larger models or thinking modes do not consistently translate into better performance. --- ## Overview  *Figure 1. Overview of the MineExplorer benchmark pipeline.* --- ## Demo Videos *Agent episode replays from MineExplorer — covering crafting, exploration, combat, and trading tasks in the Minecraft sandbox environment.* <table> <tr> <td align="center"> <video src="https://huggingface.co/datasets/jometeorie/MineExplorer-Benchmark/resolve/main/videos/craft_diamond_pickaxe.mp4" autoplay muted loop playsinline width="260"></video> <br/><sub>Craft Diamond Pickaxe</sub> </td> <td align="center"> <video src="https://huggingface.co/datasets/jometeorie/MineExplorer-Benchmark/resolve/main/videos/craft_bed_with_wool.mp4" autoplay muted loop playsinline width="260"></video> <br/><sub>Craft Bed with Wool</sub> </td> <td align="center"> <video src="https://huggingface.co/datasets/jometeorie/MineExplorer-Benchmark/resolve/main/videos/craft_a_door.mp4" autoplay muted loop playsinline width="260"></video> <br/><sub>Craft a Door</sub> </td> </tr> <tr> <td align="center"> <video src="https://huggingface.co/datasets/jometeorie/MineExplorer-Benchmark/resolve/main/videos/trade_with_villager.mp4" autoplay muted loop playsinline width="260"></video> <br/><sub>Trade with Villager</sub> </td> <td align="center"> <video src="https://huggingface.co/datasets/jometeorie/MineExplorer-Benchmark/resolve/main/videos/find_diamond_ore.mp4" autoplay muted loop playsinline width="260"></video> <br/><sub>Find Diamond Ore</sub> </td> <td align="center"> <video src="https://huggingface.co/datasets/jometeorie/MineExplorer-Benchmark/resolve/main/videos/trap_a_zombie.mp4" autoplay muted loop playsinline width="260"></video> <br/><sub>Trap a Zombie</sub> </td> </tr> <tr> <td align="center"> <video src="https://huggingface.co/datasets/jometeorie/MineExplorer-Benchmark/resolve/main/videos/reach_the_summit.mp4" autoplay muted loop playsinline width="260"></video> <br/><sub>Reach the Summit</sub> </td> <td align="center"> <video src="https://huggingface.co/datasets/jometeorie/MineExplorer-Benchmark/resolve/main/videos/defeat_spider_on_platform.mp4" autoplay muted loop playsinline width="260"></video> <br/><sub>Defeat Spider on Platform</sub> </td> <td align="center"> <video src="https://huggingface.co/datasets/jometeorie/MineExplorer-Benchmark/resolve/main/videos/cook_meat.mp4" autoplay muted loop playsinline width="260"></video> <br/><sub>Cook Meat</sub> </td> </tr> </table> --- ## Leaderboard **Table 1.** Main results on MineExplorer. P = Precision, R = Recall, A = Accuracy, MSR = Milestone Success Rate, TSR = Task Success Rate. **Bold** = best, <u>underline</u> = second best. Sorted by Overall TSR (descending). | Model | SH-P | SH-R | SH-A | SH-MSR | SH-TSR | MH-P | MH-R | MH-A | MH-MSR | MH-TSR | Ov-P | Ov-R | Ov-A | Ov-MSR | Ov-TSR | |---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---| | **Claude-Opus-4.6** | **78.70** | **74.37** | **77.58** | **77.69** | **77.69** | **59.06** | **51.79** | **57.42** | **55.04** | **23.87** | **61.91** | **54.71** | **60.33** | **58.27** | **41.08** | | Gemini-3.1-Pro-Preview | <u>74.24</u> | <u>70.35</u> | <u>73.54</u> | <u>74.23</u> | <u>74.23</u> | <u>55.76</u> | <u>49.85</u> | <u>55.05</u> | <u>52.21</u> | <u>19.53</u> | <u>58.44</u> | <u>52.50</u> | <u>57.71</u> | <u>55.36</u> | <u>37.02</u> | | Claude-Opus-4.5 | 54.77 | 50.25 | 54.14 | 54.62 | 54.62 | 50.72 | 44.93 | 48.86 | 47.21 | 14.47 | 51.31 | 45.61 | 49.62 | 48.27 | 27.31 | | GPT-5.2 | 45.84 | 45.73 | 45.46 | 45.77 | 45.77 | 43.83 | 39.25 | 42.30 | 40.92 | 9.77 | 44.12 | 40.09 | 42.76 | 41.62 | 21.28 | | GLM-5V-Turbo | 44.02 | 40.20 | 43.03 | 45.00 | 45.00 | 43.18 | 39.03 | 42.03 | 40.22 | 8.14 | 43.30 | 39.18 | 42.18 | 40.90 | 19.93 | | GPT-5.4 | 40.37 | 39.70 | 41.62 | 40.39 | 40.39 | 46.24 | 41.12 | 44.99 | 43.36 | 7.60 | 45.39 | 40.94 | 44.50 | 42.94 | 18.08 | | Claude-Sonnet-4.5 | 39.76 | 37.69 | 38.99 | 41.15 | 41.15 | 44.11 | 39.63 | 42.44 | 40.60 | 8.32 | 43.48 | 39.38 | 41.94 | 40.68 | 18.82 | | Doubao-Seed-2.0-Pro | 35.90 | 32.66 | 35.56 | 35.77 | 35.77 | 41.04 | 36.49 | 40.20 | 38.55 | 5.97 | 40.30 | 36.00 | 39.53 | 38.15 | 15.50 | | Claude-Haiku-4.5 | 35.29 | 34.17 | 34.95 | 36.54 | 36.54 | 36.25 | 32.46 | 34.69 | 33.48 | 5.61 | 36.11 | 32.68 | 34.73 | 33.92 | 15.50 | | Gemini-2.5-Flash | 37.32 | 36.68 | 36.36 | 38.08 | 38.08 | 35.01 | 31.64 | 33.78 | 34.73 | 4.70 | 35.35 | 32.29 | 34.15 | 35.24 | 15.38 | | Gemini-2.5-Pro | 36.31 | 34.67 | 35.76 | 36.54 | 36.54 | 35.67 | 32.24 | 34.39 | 32.97 | 4.16 | 35.76 | 32.55 | 34.58 | 33.48 | 14.51 | | GPT-4.1 | 28.80 | 27.64 | 29.29 | 29.62 | 29.62 | 33.98 | 30.75 | 32.48 | 31.17 | 3.98 | 33.23 | 30.34 | 32.02 | 30.95 | 12.18 | | Qwen-3-VL-235B-A22B-Instruct | 26.78 | 26.13 | 26.67 | 27.31 | 27.31 | 31.12 | 28.28 | 30.21 | 29.06 | 2.71 | 30.49 | 28.01 | 29.70 | 28.81 | 10.58 | | Qwen-3-VL-32B-Instruct | 26.17 | 27.14 | 26.47 | 26.92 | 26.92 | 28.60 | 25.75 | 27.49 | 26.68 | 2.17 | 28.25 | 25.93 | 27.34 | 26.72 | 10.09 | | LLaMA-3.2-90B-Vision-Instruct | 27.18 | 27.14 | 26.67 | 27.31 | 27.31 | 26.50 | 24.10 | 25.52 | 24.76 | 1.81 | 26.60 | 24.50 | 25.68 | 25.12 | 9.96 | | Kimi-K2.6 | 27.59 | 27.14 | 27.27 | 28.46 | 28.46 | 22.23 | 19.48 | 21.48 | 25.06 | 1.27 | 23.00 | 20.47 | 22.31 | 25.63 | 9.96 | | Qwen-3-VL-32B-Thinking | 26.37 | 25.63 | 26.67 | 26.92 | 26.92 | 28.67 | 26.05 | 27.73 | 26.88 | 1.27 | 28.34 | 25.99 | 27.57 | 26.88 | 9.47 | | Qwen-3-VL-235B-A22B-Thinking | 22.70 | 21.61 | 22.22 | 22.31 | 22.31 | 25.29 | 23.13 | 24.91 | 23.93 | 1.45 | 24.77 | 22.94 | 24.52 | 23.69 | 8.12 | SH = Single-Hop Tasks (Simple) · MH = Multi-Hop Tasks (Hard) · Ov = Overall --- ## Dataset Structure Each record (one scenario) contains: | Field | Type | Description | |---------------------|------------------|-------------| | `scene_id` | string | Zero-padded scene index, e.g. `"0000"` | | `mode` | string | Task complexity: `"single"` or `"multi"` | | `task_text` | string | Natural-language instruction shown to the agent | | `scene_name` | string | Internal name of the Minecraft scene | | `scene_description` | string | Human-readable description of the starting scene | | `commands` | list[string] | Minecraft `/commands` to set up the scene | | `selected_tasks` | list[string] | Atomic task names | | `milestones` | list[dict] | Rule-based milestone evaluators | | `reasoning_graph` | dict (nullable) | Task dependency DAG | | `design_notes` | string (nullable)| Notes from the multi-agent design workflow | ## Configs - **default** (`benchmark.jsonl`): All 813 multi-hop benchmark scenarios. - **hard** (`benchmark_hard.jsonl`): The 100 hardest scenarios subset. ## Load with 🤗 Datasets ```python from datasets import load_dataset # Full benchmark ds = load_dataset("jometeorie/MineExplorer-Benchmark", split="train") # Hard subset only ds_hard = load_dataset("jometeorie/MineExplorer-Benchmark", name="hard", split="train") ``` ## Citation ```bibtex @misc{ju2026mineexplorerevaluatingopenworldexploration, title={MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft}, author={Tianjie Ju and Yueqing Sun and Zheng Wu and Wei Zhang and Yaqi Huo and Xi Su and Qi Gu and Xunliang Cai and Gongshen Liu and Zhuosheng Zhang}, year={2026}, eprint={2605.30931}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2605.30931}, } ```



