UI-TARS
收藏资源简介:
UI-TARS是由字节跳动和清华大学联合开发的一个大规模GUI截图数据集,旨在提升GUI代理模型的感知和交互能力。该数据集包含从网站、应用程序和操作系统中收集的截图,并提取了元数据如元素类型、边界框和文本内容。数据集涵盖了元素描述、密集标注、状态转换标注、问答和标记提示等任务,支持模型在多平台上的精确交互。通过大规模动作轨迹和6M GUI教程的爬取,UI-TARS能够进行系统2推理和迭代训练,显著提升了其在动态环境中的适应性和决策能力。该数据集的应用领域包括任务自动化、工作流优化和GUI交互的智能化。
UI-TARS is a large-scale GUI screenshot dataset jointly developed by ByteDance and Tsinghua University, which is designed to improve the perception and interaction capabilities of GUI agent models. This dataset collects screenshots from websites, applications and operating systems, and extracts metadata including element types, bounding boxes and textual contents. It covers various tasks such as element description, dense annotation, state transition annotation, question answering and prompt tagging, enabling models to perform precise interactions across multiple platforms. By crawling large-scale action trajectories and 6M GUI tutorials, UI-TARS supports System 2 reasoning and iterative training, greatly enhancing its adaptability and decision-making capabilities in dynamic environments. The application scenarios of this dataset include task automation, workflow optimization and intelligent GUI interaction.




