遇见数据集

Historical SEC 8-K Market-Reaction Dataset — free sample, data dictionary and methodology

收藏
Zenodo2026-08-18 更新2026-08-20 收录
官方服务:

资源简介:

Know what each kind of company news historically did to the stock — before you build on top of it. 29,331 SEC 8-K filings from 660 US companies, read and re-classified finer than the official item codes, with each category's price reaction measured by a market-model event study: direction, magnitude, overnight-vs-session split, pre-filing baselines — and every published finding validated on two independent out-of-sample holdouts. This public repo contains the preview, data dictionary, and methodology only. The full dataset (1,668 aggregate cells + 104 validated findings, run 2026-08-14) is a one-time paid release — link and access instructions below. Historical and descriptive research data. Not investment advice, not a forecast, not a trading signal. This deposition is the free layer: a 8-row sample carrying the complete 22-column schema (preview.csv), the field-by-field data dictionary (data_dictionary.csv), and the full methodology (methodology.md). The complete dataset — 1,668 aggregate cells and 104 validated findings, run dated 2026-08-14 — is a one-time paid release at https://flinchlab.com/dataset. Contact: data@flinchlab.com. Related resources: the full methodology · every published finding, free, with holdout results · this free layer on Hugging Face (interactive table). Schema of the aggregate table fieldtypemeaningbuckettextWhich slice of events this row aggregatesbucket_typetextWhat kind of slice `bucket` isitem_codetextOfficial SEC 8-K item code, or ALL for the pooled rowitem_labeltextPlain-language name for the item codewindowtextEvent-time window in trading-day offsets, day 0 = event daywindow_typetextWhether the window ends before the filing was publicn_eventsintegerNumber of events aggregated in this cellsample_statustextSample-size policy marker. Cells under the threshold carry too_few and BLANK statistics — a blank never means 'zero effect'mean_car_pctnumberMean cumulative abnormal return over the windowmean_abs_car_pctnumberMean ABSOLUTE cumulative abnormal return — average reaction size regardless of directionmean_gap_car_pctnumberMean abnormal OVERNIGHT component (prior close to open)mean_session_car_pctnumberMean abnormal REGULAR-SESSION component (open to close)mean_abs_gap_car_pctnumberMean absolute overnight componentmean_abs_session_car_pctnumberMean absolute regular-session componentdirection_q_valuenumberBenjamini-Hochberg-adjusted q-value of the directional (BMP) test within this row's fdr_familydirection_significantbooleanWhether the directional effect survives FDR at q=0.10 within its familymagnitude_q_valuenumberBH-adjusted q-value of the reaction-size test (mean |SAR| vs noise-only null)magnitude_significantbooleanWhether excess reaction SIZE survives FDR at q=0.10 within its familyfdr_familytextMultiple-testing family this row was corrected in. q-values are NOT comparable across familiesmean_abs_scartextMean ABSOLUTE standardized CAR: each event's window CAR divided by its own forecast-error standard deviation, then averaged. Unlike mean_abs_car_pct this IS comparable across windows of different lengths. Under the no-reaction null it averages ~0.798, so a window near 0.80 is ordinary volatility however large its percentage lookscorrado_statistictextCorrado rank-test statistic for direction. A STAGGERED-SAMPLE ADAPTATION, not the textbook Corrado (1989) variance: per-event standardized rank statistics tested cross-sectionally. Documented in LIMITATIONS.mdcorrado_p_valuenumberTwo-sided p-value of the Corrado rank test. NOT FDR-corrected -- unlike direction_q_value and magnitude_q_value in this same file, which are. Provided as a distribution-free cross-check on the BMP direction result, NOT as an independent significance verdict; do not read it against a 0.05 threshold as though it were corrected

提供机构:
Zenodo
创建时间:
2026-08-17
二维码
社区交流群
二维码
科研交流群
商业服务