遇见数据集

Replication data for "RL over a Calibrated Market-Making Baseline Learns a Constant, Not a Policy"

收藏
Zenodo2026-08-19 更新2026-08-20 收录
官方服务:

资源简介:

Replication data for the working paper "RL over a Calibrated Market-Making Baseline Learns a Constant, Not a Policy" (Sanin, 2026). The paper trains a bounded-residual reinforcement-learning overlay on a calibrated Gueant-Lehalle-Fernandez-Tapia market maker over ~150 days of Bybit limit-order-book data, across five pre-registered arms that each mistune one baseline parameter. Every discriminating prediction failed, and the reference arm — required by the registration to be null — was not, routing the study to "uninterpretable" under its own committed decision rule rather than to falsification. What survives is a set of measurements about how the apparent edge arises: folding each policy's mean action on all four control dials into a single constant recovers 132% of the trained policy's edge and beats it in fourteen runs of fifteen; the policies trade 10.8% less than their baselines; and a prediction with no free parameters — each baseline's own loss rate times the volume its policy avoided — accounts for 71% of the measured edge. This deposit is the derived data behind those results, together with the verifier that re-derives every number in the manuscript from it. It exists so the paper's arithmetic can be checked by anyone, permanently, without depending on a data vendor's terms or on a code repository staying online. WHAT IS IN IT - per_day_results.csv — the 75 day-rows of the registered study (15 runs x 5 held-out days): baseline P&L, policy P&L, edge, fill counts, baseline and policy traded notional, the signed action per dial, and the untrained-control mean and dispersion.- armA_seeds.csv .. armE_seeds.csv — the 100 runs of the twenty-seed follow-up, one row per run. This is the set that showed three of five three-seed arm means inflated by 1.8x to 2.8x.- foldin/arm{A,BDE,C}.jsonl — the 100 rollouts of the four-dial static fold-in: 25 "a = 0" baselines and 75 constant-action fold-ins.- controls_armA.jsonl — the 100 rollouts of the corrected control reference: the reference arm's twenty untrained control networks re-evaluated on all five held-out days, so the reference is a distribution of policy-level means rather than of pooled control-days.- asset_b_reports/ — the 15 de-identified backtest reports behind two of the paper's tables (see below).- verify_paper_numbers.py — the verifier.- make_paper_figures.py, the manuscript (Markdown and PDF), and the submission abstract. HOW TO CHECK IT python research/verify_paper_numbers.py Python 3.10+, no dependencies. It re-derives every table cell and roughly forty prose statistics from the files above, prints one line per assertion, and exits non-zero on any disagreement. 303 assertions run from this deposit alone. The complete repository runs 333; the additional 30 are one table's AAVE rows and the tick-quantisation figures beside it, which read backtest reports that are not published. The verifier is included because the paper's central complaint is that published market-making results are not checkable. Shipping the claim without the means to test it would have repeated the thing it objects to. DE-IDENTIFICATION The paper reports a second instrument only as "Asset B" and does not name it. Its backtest reports carried the identity four ways: the symbol, absolute timestamps, one absolute price, and — decisively — the per-fill price columns of the accompanying fill and decision tables, roughly 18,000 raw prices per day. A price series of that size on a named venue can be matched back to its instrument by anyone holding the vendor's data for the month, so a ticker substitution would not have de-identified it. What is published therefore replaces the symbol with ASSET_B, re-bases every timestamp to milliseconds from its own day's start, nulls the one absolute price, renumbers the days, and withholds the per-fill tables entirely. MANIFEST.json records every field substituted. The verifier re-derives both affected tables from these de-identified files, so withholding the fills costs nothing the paper claims is checkable. WHAT IS NOT IN IT - The raw ticks. Bybit order-book and trade data from CryptoHFTData, individually retrievable from the vendor: instrument AAVEUSDT (spot and linear perpetual), training pool 2026-01-22 to 2026-06-21 end-exclusive (150 days), evaluation 2026-07-01 to 2026-07-06 end-exclusive. Redistribution is the one use such terms typically restrict, so we do not redistribute.- The simulation environment and training code, available for review on request. This deposit lets you check the analysis, not re-run the simulation.- AAVE's static k grid, one table's AAVE rows, the per-unit economics table, the live-versus-simulator calibration, the per-instrument mirror table, and the operation counts in the checklist section. These come from backtest and production reports that are not published; the verifier lists them under a heading reading "declared external — not verifiable from the released artifacts" rather than passing over them silently. LICENCE Data and documentation: CC BY 4.0. Code: MIT.

提供机构:
Zenodo
创建时间:
2026-08-19
二维码
社区交流群
二维码
科研交流群
商业服务