pandalla/datatager_standard_med_question
收藏资源简介:
--- license: apache-2.0 --- --- license: apache-2.0 --- <p align="center"> <img src="https://raw.githubusercontent.com/PandaVT/DataTager/main/assert/datatager_logo_right.png" width="650" style="margin-bottom: 0.2;"/> <p> <h5 align="center"> If you like our project, please give us a star ⭐ </h2> <h4 align="center"> [<a href="https://github.com/PandaVT/DataTager">GitHub</a> | <a href="https://datatager.com/">DataTager Home</a>] # Standard Medical Question ## Prompt for Training When training your model with this dataset, prepend the following prompt to each input instance: ``` 你需要将医疗领域中的冗长或复杂的患者咨询文本转换为简洁、结构化的问题表达。请确保输出文本保留所有关键的医疗信息,去除重复或不必要的细节,并使用专业的医疗术语准确描述患者的情况和需求。 ``` ## Description AnyTaskTune is a publication by the DataTager team. We advocate for rapid training of large models suitable for specific business scenarios through task-specific fine-tuning. We have open-sourced several datasets across various domains such as legal, medical, education, and HR, and this dataset is one of them. This dataset, titled "Standard Medical Question Data," is part of an initiative by the DataTager team under the AnyTaskTune publication. It focuses on transforming non-standard patient inquiries into standardized medical questions. This transformation aims to facilitate quicker and clearer understanding by healthcare professionals, thereby improving the efficiency of medical consultations. ## Usage This dataset is particularly valuable for training AI systems aimed at medical dialogue processing. By converting non-standard patient expressions into standardized medical queries, these AI models can assist in automating parts of the initial patient consultation process. This not only reduces the time healthcare professionals spend in understanding patient issues but also enhances the accuracy of medical advice provided. Furthermore, the dataset can be used in educational settings to train medical students on interpreting and reformulating patient questions. ## Citation Please cite this dataset in your work as follows: ``` @misc{ Extract Medical Information Dataset, author = {DataTager}, title = {Extract Medical Information Dataset}, year = {2024}, publisher = {GitHub}, journal = {GitHub repository}, howpublished = {\\url{https://github.com/PandaVT/DataTager}} } ```
许可证:Apache-2.0 --- <p align="center"> <img src="https://raw.githubusercontent.com/PandaVT/DataTager/main/assert/datatager_logo_right.png" width="650" style="margin-bottom: 0.2;"/> <p> <h5 align="center"> 如果您喜爱本项目,请为我们点亮⭐Star</h5> <h4 align="center"> [<a href="https://github.com/PandaVT/DataTager">GitHub</a> | <a href="https://datatager.com/">DataTager 官网</a>] # 标准化医疗问句 ## 训练提示词 使用本数据集训练模型时,请在每个输入样本前添加如下提示词: 你需要将医疗领域中的冗长或复杂的患者咨询文本转换为简洁、结构化的问题表达。请确保输出文本保留所有关键的医疗信息,去除重复或不必要的细节,并使用专业的医疗术语准确描述患者的情况和需求。 ## 数据集说明 AnyTaskTune 是 DataTager 团队推出的开源项目。我们倡导通过针对特定任务的微调,快速训练适配特定业务场景的大语言模型(Large Language Model)。目前我们已在法律、医疗、教育、人力资源等多个领域开源了多款数据集,本数据集即为其中之一。 本数据集命名为「标准化医疗问句数据集」,属于 DataTager 团队在 AnyTaskTune 项目框架下推出的开源计划的一部分。其核心目标是将非标准化的患者咨询文本转换为标准化医疗问句,旨在帮助医疗从业者更快速清晰地理解患者诉求,进而提升医疗问诊效率。 ## 使用场景 本数据集尤其适用于训练面向医疗对话处理的AI智能体(AI Agent)。通过将非标准化的患者表述转换为标准化医疗问句,此类AI模型可助力自动化部分初始患者问诊流程。此举不仅可减少医疗从业者理解患者诉求所花费的时间,还能提升所提供医疗建议的准确性。此外,该数据集还可用于教育场景,帮助医学生学习解读与重构患者问句。 ## 引用规范 若在研究工作中使用本数据集,请按如下格式引用: @misc{ Extract Medical Information Dataset, author = {DataTager}, title = {Extract Medical Information Dataset}, year = {2024}, publisher = {GitHub}, journal = {GitHub repository}, howpublished = {\url{https://github.com/PandaVT/DataTager}} }



