FreedomIntelligence/huatuo26M-testdatasets
收藏资源简介:
我们很高兴发布我们的评估数据集,这是Huatuo-26M的一个子集。该数据集包含6000个条目,用于我们相关研究论文中的自然语言生成(NLG)实验。我们鼓励研究人员和开发者使用此评估数据集来衡量他们自己模型的性能。这不仅是评估生成响应的准确性和相关性的机会,也是研究模型在理解和生成复杂医学语言方面能力的机会。注意:所有数据点都已匿名化,以保护患者隐私,并严格遵守数据保护和隐私法规。
We are pleased to release our evaluation dataset, which is a subset of Huatuo-26M. This dataset consists of 6000 entries and is used for natural language generation (NLG) experiments in our associated research paper. We encourage researchers and developers to utilize this evaluation dataset to benchmark the performance of their own models. This not only provides an opportunity to evaluate the accuracy and relevance of generated responses, but also to study the model's capabilities in understanding and generating complex medical language. Note: All data points have been anonymized to protect patient privacy and strictly comply with data protection and privacy regulations.
数据集概述
数据集名称
- 名称: huatuo26M-testdatasets
数据集描述
- 类别: 医学
- 语言: 中文
- 任务类别: 文本生成
- 大小: 1K<n<10K(共6,000条记录)
- 许可证: Apache-2.0
数据集详情
- 概述: 该数据集是Huatuo-26M的一个子集,包含6,000条记录,用于自然语言生成(NLG)实验。数据集旨在帮助研究人员和开发者评估其模型的性能,特别是在理解和生成复杂医学语言方面的能力。
- 隐私保护: 所有数据点均已匿名化,严格遵守数据保护和隐私法规。
引用信息
@misc{li2023huatuo26m, title={Huatuo-26M, a Large-scale Chinese Medical QA Dataset}, author={Jianquan Li and Xidong Wang and Xiangbo Wu and Zhiyi Zhang and Xiaolong Xu and Jie Fu and Prayag Tiwari and Xiang Wan and Benyou Wang}, year={2023}, eprint={2305.01526}, archivePrefix={arXiv}, primaryClass={cs.CL} }




