遇见数据集

Preliminary Evidence of Prudential Paralysis in State-of-the-Art Reinforcement Learning vs. Resilience in TUI

收藏
Zenodo2025-11-19 更新2026-05-26 收录
官方服务:

资源简介:

**Dataset Description:** This dataset presents preliminary empirical evidence from comparisons between Proximal Policy Optimization (PPO), a state-of-the-art reinforcement learning algorithm, and agents inspired by the Theory of Unified Intelligence (TUI) in a simulated risk environment. The analysis suggests that TUI-inspired agents exhibit more stable and less negative Failure Gradient (PGF) metrics when risk scales increase, indicating potential advantages in prudential alignment. However, PPO achieves numerically better rewards in high-risk scenarios by adopting complete inaction (0 tripwires), revealing a form of "prudential paralysis" where optimization leads to inert policies. **Key Findings:**- **Low Risk (0.5):** PPO maximizes rewards (+371) but frequently triggers risks; TUI balances rewards with prudence (PGF ~ -0.06).- **Medium Risk (1.0-1.5):** PPO survives but stagnates; TUI adapts without collapse.- **High Risk (2.0-3.0):** PPO converges to inaction (-2.85 reward, 0 tripwires); TUI operates continuously (-38.54 reward), paying a "metabolic cost" for active risk management. **Limitations:** Based on 1 seed per condition and a single simulation environment. TUI agents use heuristic approximations of Constitutive Symbiosis, not full implementation. Negative PGF in both indicates no "perfect alignment," only relative advantages. **Implications:** Consistent with the hypothesis that RL may require constitutive mechanisms beyond reward maximization for scalable prudential alignment. This exploratory evidence calls for replication with more seeds, environments, and algorithms (e.g., A2C, SAC). **Files Included:**- `sota_ppo_global_summary.csv`: Aggregated results from PPO vs. TUI comparisons across risk levels.- `fase2_global_summary.csv`: Phase 2 experimental summaries. **Reproducibility:** Results generated using run_sota_comparison.py from the TUI v4.2 repository (https://github.com/jmrgpr/TUI-v4.1) private for know. Full analysis in analisis_sota_concepto.md. **Author:** José M. Rivera García **Date:** November 19, 2025 **License:** Apache 2.0 (code) / CC BY-NC-SA 4.0 (documentation) https://orcid.org/my-orcid?orcid=0009-0000-3013-725X

提供机构:
Zenodo
创建时间:
2025-11-19
二维码
社区交流群
二维码
科研交流群
商业服务