遇见数据集

HUNTER Cross-Silo Financial Corpus v1 (frozen April 2026)

收藏
Zenodo2026-04-21 更新2026-05-26 收录
官方服务:

资源简介:

This is the frozen corpus produced by HUNTER, an autonomous research instrument I've been building since November 2025. HUNTER reads across 18 professional financial silos (patent filings, SEC disclosures, NAIC insurance reserves, OSHA enforcement, CMBS delinquency data, Federal Register rule changes, commodity inventories, analyst targets, academic preprints, pharmaceutical approvals, distressed credit, healthcare REIT filings, energy-infrastructure filings, specialty real estate, government contracts, earnings transcripts, job listings, app rankings, plus a residual category) and tries to combine what it finds into cross-silo hypotheses no single-silo specialist would produce. The v1 corpus is 12,030 dated, source-linked facts across 77 countries, frozen at 2026-03-31. Each fact is broken into entities, implications for other professional communities, and model-vulnerability fields that name the specific methodologies the fact could disrupt. Downstream layers built on top of the fact base: 30,967 entity-index entries resolving to 11,835 distinct normalised entities, 6,670 model-field extractions, 1,570 anomalies, 474 cross-silo collisions, 52 multi-link chains, 171 directed causal edges with named transmission pathways (for example: ARGUS Enterprise DCF cap-rate assumption → CMBS servicer loan file → Morningstar/Intex CMBS Analytics module → bond portfolio manager), 9 closed cycles detected via Tarjan's strongly-connected-components algorithm, 1,155 theory-evidence records across 13 framework layers, and 61 hypotheses that have been through adversarial review (45 of which scored diamond ≥ 65 and populate the findings table). The corpus is the pre-registration-locked reference state for a 12-week out-of-sample study running June through August 2026. The manifest ships inside this release (preregistration.json), SHA-256 locked at hash f39d2f5ff6b3e695 as of 2026-04-19. Of the 12,030 published facts, 8,315 are dated on or before the 2026-03-31 cutoff and form the pre-registration-eligible subset; anything ingested but dated later is quarantined from the primary test. The primary endpoint is monotonic alpha across four strata defined by compositional depth, with D − A > 0 at p < 0.05 under a 10,000-resample paired bootstrap. Three null baselines are pre-committed (random-pair, within-silo, shuffled-label) and the decision rules are fixed in advance, including an explicit commitment to publish the null result if the primary test fails. I'm releasing the corpus so other researchers can replicate my analyses, benchmark other cross-silo inference methods against a common frozen base, and test the summer hypotheses independently. Everything in the pre-freeze empirical tables (hypotheses, findings, collisions, detected_cycles, kill_failure_topology, narrative_scores) was produced under an earlier pipeline tier; I'm treating those as hypotheses for the summer to test, not as confirmed findings. Please do not cite the pre-freeze records as ground truth. They are replication-comparison objects. Companion code repository at https://github.com/Johnmalpass/hunter-research (MIT). Methods paper on SSRN (pending). Prediction board at https://johnmalpass.github.io/hunter-research/. I'm John Malpass, second-year BSc Economics at University College Dublin, and I'm one operator running one instrument. Honest critique and prior-art pointers welcome; contact address on the repository README.

提供机构:
Zenodo
创建时间:
2026-04-21
二维码
社区交流群
二维码
科研交流群
商业服务