osunlp/ACuRL
收藏资源简介:
该数据集包含由ACuRL(自主课程强化学习框架)生成的课程任务,用于使计算机使用代理能够持续适应目标环境,无需人类数据。数据集分为两个部分:qwen3vl和uitars,分别对应两个基础代理:Qwen3-VL-8B-Instruct和UI-TARS-1.5-7B。每个部分包含4,608个自然语言任务,每个任务针对特定环境和课程迭代生成。数据源涵盖多个环境,如LibreOffice Impress、Calc、Writer、Thunderbird、Celestia和KAlgebra,每个数据源包含768个任务,组织为3个课程迭代(迭代1、2、3),每个迭代256个任务。数据集列包括data source(数据源,表示环境)、iteration(迭代,表示课程迭代编号)和task(任务,表示生成的自然语言任务)。
This dataset contains curriculum tasks generated by ACuRL, an Autonomous Curriculum Reinforcement Learning framework for continually adapting computer-use agents to target environments with zero human data. The dataset includes two splits, qwen3vl and uitars, corresponding to two base agents: Qwen3-VL-8B-Instruct and UI-TARS-1.5-7B. Each split contains 4,608 natural-language tasks generated for specific environments and curriculum iterations. Data sources include environments such as LibreOffice Impress, Calc, Writer, Thunderbird, Celestia, and KAlgebra, with each data source containing 768 tasks organized into 3 curriculum iterations (iteration 1, 2, 3), each with 256 tasks. Dataset columns are data source (indicating the environment), iteration (indicating the curriculum iteration number), and task (indicating the generated natural-language task).




