pit-earnings-call-qa
收藏资源简介:
该数据集是用于PIT(Point-in-Time)系列语言模型监督微调的金融问答数据集,专门针对美国上市公司财报电话会议记录构建。数据集旨在微调Diamegs/PIT-4B-FT-*模型快照,同时严格遵守PIT时间顺序原则——训练中不使用基础模型知识截止日期之后的任何会议记录。数据集包含两个时间快照(202112和202212),每个快照按时间顺序划分为训练集、验证集、测试集和基准集。数据规模在10万到100万样本之间,具体到202212快照包含189,362个训练样本、16,778个验证样本、16,900个测试样本和1,000个基准样本。数据内容涵盖四种问答类型:基于LLM生成问题的正向合成问答、基于实际分析师问题的正向自然问答、给定管理层回答的反向问题生成,以及无法回答问题的识别训练。数据以两种格式提供:扁平记录格式(包含transcript_id、context、question、answer等字段)和聊天格式(适合训练脚本使用)。该数据集专门用于金融领域的问答任务,特别是财报电话会议内容的理解与分析。
This dataset is a financial question-answering dataset for supervised fine-tuning of the PIT (Point-in-Time) series of language models, specifically constructed from earnings call transcripts of U.S. publicly traded companies. It is designed to fine-tune the Diamegs/PIT-4B-FT-* model snapshots while strictly adhering to the PIT chronological principle—no meeting records after the base models knowledge cutoff date are used in training. The dataset includes two time snapshots (202112 and 202212), each chronologically divided into training, validation, test, and benchmark sets. The data scale ranges from 100,000 to 1,000,000 samples, with the 202212 snapshot specifically containing 189,362 training samples, 16,778 validation samples, 16,900 test samples, and 1,000 benchmark samples. The data content covers four types of question-answering: forward synthetic QA based on LLM-generated questions, forward natural QA based on actual analyst questions, reverse question generation given management responses, and training for identifying unanswerable questions. The data is provided in two formats: a flat record format (including fields such as transcript_id, context, question, answer) and a chat format (suitable for training scripts). This dataset is specifically designed for financial question-answering tasks, particularly for understanding and analyzing earnings call content.




