EarningsCall_Dataset
收藏资源简介:
这是一个包含S&P 500公司盈利电话会议的数据集,包括文本记录和相应的音频记录。数据集用于研究CEO的言语和声音特征如何影响金融风险预测。
This dataset comprises earnings call transcripts and corresponding audio recordings from S&P 500 companies. It is utilized to investigate how CEOs' speech and vocal characteristics influence financial risk prediction.
数据集概述
数据集名称
What You Say and How You Say It Matters: Predicting Financial Risk Using Verbal and Vocal Cues
数据集内容
- 数据类型:包含文本和音频数据。
- 数据来源:S&P 500公司的盈利电话会议记录。
- 数据描述:每个文件夹代表一次盈利电话会议,文件夹名为“公司名_日期”。每个文件夹内包含处理过的文本记录和分割的音频记录。文本记录中每行代表CEO的一句话,按时间顺序排列。音频记录按“发言人_段落_句子”命名,通过迭代强制对齐(IFA)处理,首先按段落分割,再按句子分割。
数据集用途
用于预测金融风险,通过分析CEO在盈利电话会议中的言语(文本)和声音(音频)特征。
数据集访问
- 完整数据集:由于数据集体积较大,无法在GitHub上完整存储,已上传至Google Drive,可通过提供的链接下载。
- 示例数据:GitHub仓库中包含少量示例数据。
引用信息
-
作者:Yu Qin 和 Yi Yang
-
发表年份:2019年
-
发表会议:第57届计算语言学年会
-
引用格式:
@InProceedings{P19-xxxx, author = "Qin, Yu and Yang, Yi", title = "What You Say and How You Say It Matters: Predicting Financial Risk Using Verbal and Vocal Cues", booktitle = "Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)", year = "2019", publisher = "Association for Computational Linguistics", pages = "xx--xx", location = "Florence , Italy", url = "" }




