遇见数据集

Friction-MARL: Factorial Results and Control Battery for the Consent-Friction Functional in Multi-Agent Coordination

收藏
Zenodo2026-06-27 更新2026-06-12 收录
官方服务:

资源简介:

Code and complete results for the Friction-MARL study — the empirical companion to The Axiom of Consent (arXiv:2601.06692), testing the consent-friction functional F = σ(1+ε)/(1+α) among Independent Q-Learning (IQL) and Value-Decomposition (VDN) agents in a shared-resource coordination MDP (four agents, three capped continuous resources; reward is the negative weighted squared distance from each agent's ideal state). Version 2.0 adds a control battery to the original 5×5×5 factorial: σ-normalization: the apparent “stakes dominate” effect is mechanical (degree-1 reward scaling); on the stake-normalized gap, preference alignment is the dominant structural factor. Value decomposition (VDN): the coordination patterns are not specific to independent learning. Signed preference DGP: the original factorial's α varied only the magnitude of an always-positive cross-agent correlation; a genuinely signed equicorrelated design (plus two-team and n=2 designs reaching ρ=−1) shows the earlier symmetric “U-shape” was a sign-blind data-generating-process artifact. Separable-resource control: the surviving cooperation effect flattens when agents act on independent pools — the friction is shared-state contention, not preference correlation per se (bounded to common-pool environments). n=2 clean strong-opposition: at genuine unconfounded ρ=−1, opposition does not beat indifference; only cooperative alignment reduces coordination friction. Contents: the friction_marl package (environment, IQL/VDN agents, factorial design), run scripts, and results/ for the full factorial (CPU and GPU), GPU/CPU cross-validation, and the control experiments (review_upgrade, adversarial_rerun including the functional-form refit, reviewer3_controls). Per-experiment FINDINGS.md files and the accompanying paper are authoritative for interpretation. Reproducibility: Python; see requirements.txt / pyproject.toml. Heavy raw artifacts and figure renders that are excluded from the public GitHub repository are included here. License CC-BY-4.0.

本数据集包含一项5×5×5析因多智能体强化学习(Multi-Agent Reinforcement Learning, MARL)研究的完整结果,该研究聚焦4个独立Q学习(Independent Q-Learning, IQL)智能体对3个共享、带容量限制的连续型资源的协调问题。奖励函数定义为各智能体实际资源状态与理想资源状态的加权平方距离的负值。 该析因实验包含3个参数,每个参数设置5个水平,每种实验条件下开展30次独立重复实验,每次重复包含1000个智能体交互回合,共计125种实验条件: 1. α(取值:−0.8、−0.4、0.0、0.4、0.8):偏好对齐度,即智能体目标偏好的相关结构(合作型/无关型/对抗型)。该参数表征分歧的结构,而非摩擦强度。 2. σ(取值范围:0.2–1.0):偏好强度/利害权重,即智能体对各资源的关注程度。 3. ε(取值范围:0.0–1.0):观测噪声,即智能体对资源状态观测时引入的高斯噪声。 本研究采用两种独立实现方案(CPU多工作进程与GPU向量化实现)开展实验,并进行了交叉验证(斯皮尔曼秩相关系数ρ=0.937,排名一致性良好)。核心研究结论如下:α参数呈现U型分布——结构化分歧(合作或对抗型偏好)优于无关型偏好;利害权重的影响优于对齐结构;摩擦可均衡各智能体的实验结果;奖励函数收敛无需策略收敛;观测噪声对实验结果几乎无影响。 数据集包含以下内容:各实验条件与各重复实验的CSV文件、各重复实验的学习曲线与最终策略向量(.npz格式)、完整的统计分析套件(包括热图、残差/模型比较图、方差分析与回归表格、动力学/聚类可视化图)、GPU-CPU交叉验证结果、可直接用于学术论文的LaTeX表格、分析报告,以及运行与分析代码(friction_marl工具包)。完整的目录结构与列模式请参见README.md文件。 本数据集为配套的多智能体强化学习论文(DAI-2606,《当利害权重占据主导地位》,即将刊出)提供了数据来源,并为《同意公理》(arXiv:2601.06692)与《复制子优化机制》(arXiv:2601.06363)中的摩擦算子提供了实证基础。

提供机构:
Zenodo
创建时间:
2026-06-11
二维码
社区交流群
二维码
科研交流群
商业服务