BBT-FinCorpus
收藏资源简介:
BBT-FinCorpus是由上海数据科学重点实验室创建的大型中文金融领域数据集,包含约300GB的原始文本,来源于金融新闻、公司公告、研究报告和社交媒体等四个不同渠道。该数据集的创建旨在丰富金融领域的文本多样性,支持金融预训练语言模型的开发。通过精细的收集和处理,BBT-FinCorpus覆盖了金融NLP任务中的主要文本类型,为金融领域的语言理解和生成任务提供了丰富的数据资源。该数据集的应用领域广泛,特别适用于金融信息提取、情感分析等任务,旨在提升中文金融NLP的整体水平。
The BBT-FinCorpus is a large-scale Chinese financial domain dataset created by Fudan University, containing approximately 300GB of raw text sourced from four different channels, including financial news, company announcements, research reports, and social media. The establishment of this dataset aims to enrich the diversity of text in the financial field and support the development of financial pre-trained language models. Through meticulous collection and processing, the BBT-FinCorpus covers the main text types in financial NLP tasks, providing a rich data resource for language understanding and generation tasks in the financial field. The dataset has a wide range of applications and is particularly suitable for tasks such as financial information extraction and sentiment analysis, with the goal of enhancing the overall level of Chinese financial NLP.

- 1BBT-Fin: Comprehensive Construction of Chinese Financial Domain Pre-trained Language Model, Corpus and Benchmark上海数据科学重点实验室 · 2023年



