Nemotron-SpecializedDomains-Finance-v1
收藏资源简介:
Nemotron-SpecializedDomains-Finance 是一个大规模合成的金融问答数据集,旨在提升大型语言模型在专业金融推理和文档理解任务上的表现。该数据集包含超过326,000个高质量的问答对,这些问答对基于2019年至2024年间标普500公司的SEC文件生成。数据集采用模板化合成数据生成(SDG)方法,确保所有问题和答案都锚定在SEC文件的具体章节上,保证了事实准确性。数据集覆盖公司财务、风险因素、财务表现、治理、合规及业务运营等多个领域,并通过GenSelect方法过滤,确保回答的连贯性、准确性和上下文相关性。数据集格式为JSONL,每条样本包含角色化的对话结构(系统、用户、助手消息)和元数据。适用于金融领域专家系统的监督微调、特定领域推理、文档理解及金融问答系统开发等场景。数据集遵循CC BY 4.0许可,适合商业用途。
Nemotron-SpecializedDomains-Finance is a large-scale synthetic financial question-answering dataset aimed at enhancing the performance of large language models on professional financial reasoning and document understanding tasks. It contains over 326,000 high-quality question-answer pairs generated based on the SEC filings of S&P 500 companies from 2019 to 2024. The dataset employs the templated Synthetic Data Generation (SDG) methodology, ensuring that all questions and answers are anchored to specific sections of the SEC filings, thereby guaranteeing factual accuracy. It covers a wide range of domains including corporate finance, risk factors, financial performance, corporate governance, compliance, and business operations. It is filtered using the GenSelect method to ensure the coherence, factual accuracy, and contextual relevance of the answers. The dataset is stored in JSONL format, with each sample containing a role-based dialogue structure (system, user, and assistant messages) along with metadata. It is applicable to scenarios such as supervised fine-tuning for financial domain expert systems, domain-specific reasoning, document understanding, and the development of financial question-answering systems. The dataset is licensed under CC BY 4.0 and is suitable for commercial use.



