Multi-Channel Sentiment Corpus for US (S&P 500) and Indian (Nifty 500) Equity Markets: Speaker-Role-Segmented Earnings-Call Transcripts and News Headlines
收藏资源简介:
This deposit provides a reusable sample and the complete construction pipeline for a dual-channel financial sentiment corpus covering the constituents of the S&P 500 (United States) and Nifty 500 (India) indices. Channel A comprises FinBERT-scored financial news headlines (48,361 US; 44,640 India). Channel B comprises earnings-call transcript chunks (326,749 US; 291,992 India), each labelled with speaker role (management, analyst, moderator) with a per-label confidence tier and a call-phase tag (prepared remarks vs. question-and-answer), yielding a three-way segmentation not available in prior public earnings-call datasets. To protect ongoing derivative research and respect the redistribution terms of the underlying exchange filings, this public record contains representative 2,000-row samples of each dataset, the complete processing pipeline (allowing regeneration of the full corpus from public source filings), hand-labelled validation samples, a full data dictionary, and aggregate statistics. The complete scored corpus is available from the corresponding author on reasonable request. Transcripts were scored with yiyanghkust/finbert-tone and news with ProsusAI/finbert; both models were selected empirically against hand-labelled ground truth. The dataset supports research in financial sentiment analysis, cross-market comparison, disclosure-tone analysis, and speaker-role-conditional language modelling.



