shuaishuaicdp/GUI-World
收藏资源简介:
GUI-World数据集引入了一个全面的基准,用于评估多模态大语言模型(MLLMs)在动态和复杂的图形用户界面(GUI)环境中的表现。它包含了六种GUI场景和八种GUI导向的问题类型的广泛注释,评估了当前最先进的图像和视频大语言模型在处理动态和多步骤任务时的局限性。该数据集为未来研究提供了宝贵的见解和基础,旨在增强MLLMs对动态GUI内容的理解和交互能力,推动能够感知和交互静态及动态GUI元素的强大GUI代理的发展。
GUI-World introduces a comprehensive benchmark for evaluating MLLMs in dynamic and complex GUI environments. It features extensive annotations covering six GUI scenarios and eight types of GUI-oriented questions. The dataset assesses state-of-the-art ImageLLMs and VideoLLMs, highlighting their limitations in handling dynamic and multi-step tasks. It provides valuable insights and a foundation for future research in enhancing the understanding and interaction capabilities of MLLMs with dynamic GUI content. This dataset aims to advance the development of robust GUI agents capable of perceiving and interacting with both static and dynamic GUI elements.
数据集:GUI-World
概述
GUI-World 引入了一个全面的基准,用于评估多模态大型语言模型(MLLMs)在动态和复杂图形用户界面(GUI)环境中的表现。该数据集包含六个 GUI 场景和八种类型的 GUI 导向问题的广泛注释。它评估了最先进的图像大型语言模型(ImageLLMs)和视频大型语言模型(VideoLLMs),突出了它们在处理动态和多步骤任务方面的局限性。GUI-World 提供了宝贵的见解,并为未来在增强 MLLMs 对动态 GUI 内容的理解和交互能力方面的研究奠定了基础。该数据集旨在推动开发能够感知和与静态和动态 GUI 元素交互的强大 GUI 代理的发展。
数据集信息
- 任务类别: 问答、文本生成
- 语言: 英语
- 数据集大小: 10K<n<100K
- 许可证: Creative Commons Attribution 4.0 International License
引用
@article{chen2024gui, title={GUI-WORLD: A Dataset for GUI-Orientated Multimodal Large Language Models}, author={GUI-World Team}, year={2024} }




