遇见数据集

TRIATT-CAR: A leakage-controlled and uncertainty-aware machine-learning framework for metadata-based influence scoring

收藏
Zenodo2026-08-04 更新2026-08-13 收录
官方服务:

资源简介:

Replication package for a study that ranks financial analysts by the out-of-sample predictive power of their published commentary, using a subject-relation-object attention model with a deep-ensemble uncertainty decomposition, validated against cumulative abnormal returns (CAR). CONTENTS - Complete pipeline as runnable Python scripts and as the original Jupyter notebooks (outputs stripped, credentials removed): Elasticsearch retrieval, cleaning and record linkage, KeyBERT keyword extraction with BIO tagging, a four-way keyword-extractor comparison, feature construction, model training and every reported statistical test. - De-identified, standardised model input tables: 13,324 analyst-article records covering 2,346 analysts and 7 companies (META, GOOG, MSFT, AAPL, AMZN, TSM, NVDA), published 2005-10-21 to 2026-03-16, with all article text removed. - Daily split- and dividend-adjusted closing prices for the six modelled tickers plus SPY (5,149 trading days), so the abnormal-return target is reproducible offline rather than depending on a live data feed. - Per-document scores behind the keyword-extractor comparison: 19,784 documents x 4 extractors (KeyBERT, YAKE, TextRank, TF-IDF) x 4 reference-free metrics, plus Friedman omnibus tests, Kendall's W effect sizes and Nemenyi post-hoc matrices. - All analyst-level and article-level model outputs, ranking baselines, predictive-validity tests and figures, for both the reported market-model CAR specification and the market-adjusted robustness specification. - A clearly labelled synthetic example file that allows the licence-restricted stages of the pipeline to be executed without a Seeking Alpha licence. RESTRICTED CONTENT The raw corpus consists of analyst articles and profile pages from Seeking Alpha, a commercial platform licensed to the authors' institution under terms that do not permit redistribution. Article titles, summaries and full texts, and analyst display names, are therefore not included. This restriction applies to raw text only and does not restrict the quantitative material the study's findings rest on: every reported number, table and figure can be regenerated from the files supplied here. Researchers holding an institutional Seeking Alpha licence can regenerate the restricted intermediates with the retrieval and linkage scripts provided. Analysts appear only as integer identifiers; the mapping to display names is not distributed. See README.md and DATA_AVAILABILITY.md in the package for full documentation, including a column-level data dictionary and the equivalence checks demonstrating that removing the text columns leaves every model feature unchanged.

提供机构:
Zenodo
创建时间:
2026-08-04
二维码
社区交流群
二维码
科研交流群
商业服务