遇见数据集

Replication Package for "The Limits of Reinforcement Learning in Statistical Arbitrage: A Large-Universe Benchmark Against Classical Mean Reversion"

收藏
Zenodo2026-06-24 更新2026-06-28 收录
官方服务:

资源简介:

RL_StatArb_Replication_Russell3000 is the public replication package for the paper: “The Limits of Reinforcement Learning in Statistical Arbitrage: A Large-Universe Benchmark Against Classical Mean Reversion.” This package provides the public-safe replication materials for a large-universe statistical-arbitrage study comparing a transparent classical z-score mean-reversion benchmark with four reinforcement-learning agents: Proximal Policy Optimization, Advantage Actor-Critic, Soft Actor-Critic, and Twin Delayed Deep Deterministic Policy Gradient. The study uses a Bloomberg Russell 3000-based equity environment, dynamic pair selection, transaction costs, market-regime variables, and rolling walk-forward validation over 2016–2025. The empirical design compares discrete-action agents, PPO and A2C, with continuous-action agents, SAC and TD3, while also testing a reward-shaped A2C extension designed to reduce excessive exposure. The Russell 3000 data pipeline began with 2,901 exported Bloomberg members, successfully downloaded OHLCV and market-capitalization data for 2,892 securities, and retained 1,269 securities in the filtered tradable universe. The results show that greater algorithmic flexibility does not automatically improve statistical-arbitrage performance. The z-score benchmark achieves the strongest average risk-adjusted results, producing the highest Sharpe, Sortino, and Calmar ratios across the walk-forward test windows. PPO and A2C learn active trading policies but remain overexposed and fail to outperform the benchmark after transaction costs. Directional reward shaping reduces A2C’s active rate from 96.1% to 33.5% and improves drawdown control, but it does not close the risk-adjusted performance gap. SAC and TD3 also underperform, despite their ability to adjust position size continuously. The package therefore supports replication of the paper’s central finding: standard reinforcement-learning agents require stronger portfolio-level risk control before they can reliably improve on transparent mean-reversion rules. Full Public-Safe Code Pipeline Included This replication package includes the full public-safe code pipeline from Bloomberg data collection to model training, comparison, and statistical validation: 01 Build Russell 3000 ticker list from Bloomberg export02 Download Bloomberg OHLCV and market-cap data03 Download market-regime data04 Download Bloomberg metadata05 Clean/filter stock panel and build tradable universe06 Dynamic pair selection07 Build RL pair dataset08 Train PPO and A2C09 Train directional reward-shaped A2C10 Train SAC continuous-action model11 Train TD3 continuous-action model12 Compare all algorithms13 Run statistical tests The code implements a common empirical design so that all models are evaluated using the same data environment, dynamic pair-selection process, transaction-cost assumption, and walk-forward validation protocol. The methodology compares four model families: the classical z-score benchmark, discrete-action PPO/A2C agents, reward-shaped A2C, and continuous-action SAC/TD3 agents. Bloomberg Data Restriction Notice This study uses Bloomberg-derived data, including: Russell 3000 equity OHLCV dataMarket capitalization dataSector and industry classification dataS&P 500 market dataVIX dataU.S. Treasury yield dataBloomberg Russell 3000 membership information Due to Bloomberg and institutional data-licensing restrictions, the raw and processed Bloomberg-derived datasets are not redistributed in this public Zenodo record. This package does not include: Raw Bloomberg OHLCV filesProcessed Parquet datasetsRussell 3000 membership exportsBloomberg GICS metadata exportsMarket-regime data files derived from BloombergTrained models derived from Bloomberg dataDetailed trade logsDaily return logs Researchers with appropriate Bloomberg access may reproduce the dataset by running the provided data-collection and preprocessing scripts in the documented order. Package Contents This public replication package includes: Source codeRun-order documentationData-availability statementBloomberg data-restriction noticeAggregate empirical resultsAll-algorithm comparison tablesStatistical-test summariesAPA-style manuscript tablesColor figuresFigure source dataReplication metadataCitation fileRequirements file It is designed to support transparency and reproducibility while respecting Bloomberg data-licensing restrictions. Creator Metadata Veliota DrakopoulouHigher Colleges of Technology, United Arab EmiratesEmbry-Riddle Aeronautical University, United StatesORCID: 0000-0002-1670-8033Contact: vdrakopoulou@hct.ac.ae Suggested Keywords Statistical arbitragePairs tradingReinforcement learningMean reversionZ-score tradingPPOA2CSACTD3Reward shapingPosition sizingCointegrationRussell 3000BloombergOHLCVWalk-forward validationAlgorithmic tradingIntelligent trading systemsQuantitative finance

提供机构:
Zenodo
创建时间:
2026-06-24
二维码
社区交流群
二维码
科研交流群
商业服务