遇见数据集

GameQA-140K

收藏
魔搭社区2026-07-15 更新2026-07-15 收录
官方服务:

资源简介:

## 🎊 News * [2026/07] 🔥**Peking University and WeChat AI** use our Game-RL data synthesis code in the ICML 2026 paper [ReaForest](https://openreview.net/forum?id=Ktvb8rqFV7) for 2D Maze, 3D Maze, and Sokoban spatial-planning tasks. * [2026/06] 🔥**Shanghai AI Lab** uses our Game-RL codebase in [MoTiF](https://arxiv.org/abs/2606.12886), building 15K+ Sokoban and Maze interleaved-thinking samples. * [2026/04] 🔥**Princeton University** uses our GameQA dataset in their [Vero](https://github.com/zlab-princeton/vero) project. * [2026/03] 🔥**National University of Singapore** uses our games in the [Gym-V](https://arxiv.org/pdf/2603.15432) platform. * [2026/02] 🔥**Alibaba Group and Shanghai Jiao Tong University** use our GameQA-140K dataset at scale in the [DeepVision-103K](https://huggingface.co/datasets/skylenage/DeepVision-103K#%F0%9F%99%8F-acknowledgements) dataset, which accounts for around 50% of its "visual logic problems". * [2026/01] 🔥**Shanghai AI Lab** uses our GameQA-140K dataset at scale in the [MMFineReason](https://mmfinereason.github.io/) dataset, which accounts for **87.65%** of its "Puzzle/Game" samples. * [2026/01] 🔥**THUML and ByteDance Seed** use our Sokoban code for the synthesis of the Sokoban task samples in [VisWorld-Eval](https://thuml.github.io/Reasoning-Visual-World/) (and the training data). * [2026/01] 🔥🔥*Our work has been accepted by* **ICLR 2026**! 🎉🎉🎉 * [2025/11] 🔥**DeepWisdom and Tsinghua University** use the maze-like games in our GameQA dataset in the [VR-Bench](https://github.com/FoundationAgents/VR-Bench) benchmark, which evaluates video models' reasoning. * [2025/11] 🔥**Shanghai Innovation Institute** uses the games in our GameQA dataset for image editing reasoning tasks ("game-world scenarios"), developing the [UniREditBench](https://maplebb.github.io/UniREditBench/) benchmark and the [UniREdit-Data-100K](https://huggingface.co/datasets/maplebb/UniREdit-Data-100K) training data. ## 1. Overview GameQA is a large-scale, diverse, and challenging multimodal reasoning dataset designed to enhance the general reasoning capabilities of Vision Language Models (VLMs). Generated using the innovative Code2Logic framework, it leverages game code to synthesize high-quality visual-language Chain-of-Thought (CoT) data. The dataset addresses the scarcity of multimodal reasoning data, critical for advancing complex multi-step reasoning in VLMs. Each sample includes visual game state, targeted question, original analysis, augmented step-by-step reasoning (`refinement`) and final answer, derived from the logical structures inherent in game code. 📖 Paper: [Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs’ General Reasoning](https://arxiv.org/abs/2505.13886) 🔗 Website: https://iclr26-game-rl.github.io 💻 Code: https://github.com/tongjingqi/Game-RL ## 2. Dataset Files For a quick preview of the dataset, `GameQA-data_studio-preview.parquet` contains **300 sampled entries from the training set**. This file is optimized for online viewing in tools like Data Studio and includes embedded image data. For the **full training dataset**, please download `games_data.json`. For the **full test dataset**, please download `games_data_test.json`. Associated image files are available in `games_images.zip` and `games_images_test.zip` respectively. ## 3. Dataset Description | **Attribute** | **Description** | |------------------------------|-----------------------------------------------------------------------------------------------------| | **Size** | ~140,000 question-answer pairs (126,760 training, 15,047 testing). | | **Diversity** | 30 unique games, 158 distinct tasks covering various cognitive skills. | | **Game Categories** | - 3D Spatial Perception and Understanding<br>- Pattern Recognition and Matching<br>- Multi-step Reasoning<br>- Strategic Planning | | **Format** | Visual Question Answering (VQA):<br>- Game state image<br>- Targeted question<br>- Step-by-step reasoning<br>- Final answer | | **Question Types** | - Multiple-choice (typically 7-8 options)<br>- Fill-in-the-blank (e.g., numbers, coordinates) | | **Challenging** | Difficult for SOTA VLMs (<50% accuracy on test set). | | **Scalability & Cost** | Code2Logic enables massive-scale generation with minimal cost after initial setup. | | **Difficulty Levels** | - **Plot Level (Image Complexity):** Easy, Medium, Hard<br>- **QA Level (Task Complexity):** Easy, Medium, Hard | ## 🔎 Citation If you find our work or dataset useful, we would appreciate it if you could cite our work: ```bibtex @article{tong2025game, title={Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning}, author={Tong, Jingqi and Tang, Jixin and Li, Hangcheng and Mou, Yurong and Zhang, Ming and Zhao, Jun and Wen, Yanbo and Song, Fan and Zhan, Jiahao and Lu, Yuyang and others}, journal={arXiv preprint arXiv:2505.13886}, year={2025} } ```

提供机构:
maas
创建时间:
2026-01-31
二维码
社区交流群
二维码
科研交流群
商业服务