drug-dev-sft-dataset
收藏资源简介:
本数据集是一个专为药物研发领域大语言模型监督微调(SFT)而构建的中文数据集。它旨在涵盖从药物发现到上市后监测的完整药物研发流程知识,为模型提供该垂直领域的专业训练数据。数据集包含7174条高质量的问答对,总token数超过2065万。数据采用Alpaca JSON格式,每条数据由`instruction`(问题/指令)和`output`(回答)两个字段构成。数据内容覆盖了药物研发的15个核心专业领域,包括药物靶点发现与验证、药物化学与分子设计、药物分析方法、药物安全性评价、药物代谢与药代动力学、临床试验各阶段要点、国际多中心临床试验、药物注册申请流程、药物一致性评价、药物警戒与不良反应监测、药物联合治疗策略、创新药物研发案例、生物制品研发、中药现代化以及药物研发流程概述。数据通过DeepSeek-chat模型生成,并经过严格的人工审核抽样,质量评分达到4.5-5.0的优秀标准,审核保留率为100%,确保了数据的高质量和可靠性。该数据集适用于训练或微调面向药物研发专业任务的大语言模型。
This dataset is a Chinese dataset specifically constructed for supervised fine-tuning (SFT) of large language models in the drug research and development field. It aims to cover the complete drug R&D process knowledge from drug discovery to post-marketing surveillance, providing professional training data for models in this vertical domain. The dataset contains 7174 high-quality question-answer pairs, with a total token count exceeding 20.65 million. The data is in Alpaca JSON format, with each entry consisting of two fields: `instruction` (question/instruction) and `output` (answer). The content covers 15 core professional areas of drug R&D, including drug target discovery and validation, medicinal chemistry and molecular design, drug analysis methods, drug safety evaluation, drug metabolism and pharmacokinetics, key points of clinical trial phases, international multi-center clinical trials, drug registration application processes, drug consistency evaluation, pharmacovigilance and adverse reaction monitoring, drug combination therapy strategies, innovative drug R&D case studies, biological product development, modernization of traditional Chinese medicine, and an overview of the drug R&D process. The data was generated using the DeepSeek-chat model and underwent rigorous manual audit sampling, achieving an excellent quality score of 4.5-5.0, with an audit retention rate of 100%, ensuring high quality and reliability. This dataset is suitable for training or fine-tuning large language models for professional drug R&D tasks.




