Finance-Instruct-500k
收藏资源简介:
Finance-Instruct-500k数据集是一个专门为金融任务、推理和多轮对话训练高级语言模型而设计的综合性数据集。该数据集包含了超过500,000条条目,涵盖了金融问答、推理、情感分析、主题分类、多语言命名实体识别(NER)和对话AI等多个方面。数据集的特点包括广泛的数据覆盖、多轮对话、多样化的数据来源、RAG格式的数据、去重和预处理、以及XBRL标记等。数据集支持多种任务和用例,如金融问答、推理任务、对话AI、NER、情感分析、主题分类、轻量级LLM训练和RAG应用。数据集的结构包括系统、用户和助手字段,主要语言为英语和中文。数据集的收集和预处理过程包括去重、数据清洗、数据集合并、注释和XBRL标记。此外,数据集还考虑了用户隐私和伦理问题,并指出了数据集的局限性和引用方式。
Finance-Instruct-500k is a comprehensive dataset designed for training advanced language models on financial tasks, reasoning, and multi-turn conversations. This dataset contains over 500,000 entries covering a wide range of areas including financial question answering, reasoning, sentiment analysis, topic classification, multilingual named entity recognition (NER), and conversational AI. Key features of this dataset include extensive data coverage, multi-turn dialogues, diverse data sources, RAG-formatted data, deduplication and preprocessing, as well as XBRL tagging. The dataset supports multiple tasks and use cases, such as financial question answering, reasoning tasks, conversational AI, NER, sentiment analysis, topic classification, lightweight LLM training, and RAG applications. The structure of the dataset includes system, user, and assistant fields, with the primary languages being English and Chinese. The collection and preprocessing pipeline of the dataset includes deduplication, data cleaning, dataset merging, annotation, and XBRL tagging. In addition, the dataset takes into account user privacy and ethical considerations, and specifies the limitations and citation guidelines of the dataset.




