jwjeonn/smacv2-offline
收藏资源简介:
SMACv2离线多智能体强化学习数据集是基于SMACv2(星际争霸多智能体挑战赛v2)基准收集的离线强化学习数据集。数据通过QMIX策略生成,以HDF5格式存储为episode批次,适用于离线多智能体强化学习研究。数据集包含三个种族(Protoss、Terran、Zerg)和三种团队规模(3v3、5v5、10v10)的场景,每个场景提供两种质量的数据集:medium(来自中等水平策略的轨迹)和medium-replay(训练到中等水平过程中积累的回放缓冲区,包含不同质量的行为混合)。文件格式包括观察值、状态、动作、奖励等关键数据,形状和维度因场景而异。使用HuggingFace Hub可下载和使用数据集,许可证为CC-BY-4.0。
SMACv2 Offline Multi-Agent RL Datasets are offline reinforcement-learning datasets collected on the SMACv2 (StarCraft Multi-Agent Challenge v2) benchmark. Trajectories were generated by QMIX policies and stored as episode batches in HDF5 format, suitable for offline MARL research. The datasets include scenarios for three races (Protoss, Terran, Zerg) and three team sizes (3v3, 5v5, 10v10), with two dataset qualities provided for each scenario: medium (trajectories from a mid-level policy) and medium-replay (the replay buffer accumulated during training up to the medium level, containing a mixture of behaviors of varying quality). The file format includes keys such as observations, state, actions, and reward, with shapes and dimensions varying by scenario. The datasets are accessible via HuggingFace Hub and released under the CC-BY-4.0 license.




