MeDAL
收藏资源简介:
MeDAL是由麦吉尔大学创建的一个大型医学文本数据集,专注于医学缩写消歧,旨在支持医学领域自然语言理解的预训练。该数据集包含14,393,619篇文章,平均每篇文章包含3个缩写。数据集的创建过程利用了PubMed摘要,通过逆向替换技术生成样本,无需人工标注。MeDAL数据集的应用领域广泛,主要用于提高模型在医学文本处理中的性能,特别是在缩写消歧任务上,有助于提升模型在下游医学任务中的表现和收敛速度。
MeDAL is a large-scale medical text dataset developed by McGill University, which focuses on medical abbreviation disambiguation and aims to support pre-training for natural language understanding in the medical domain. This dataset contains 14,393,619 articles, with an average of 3 abbreviations per article. The dataset was constructed using PubMed abstracts, and samples were generated via reverse substitution technology without requiring manual annotation. MeDAL has a wide range of application scenarios, mainly used to improve the performance of models in medical text processing, especially in abbreviation disambiguation tasks, and it helps to enhance the performance and convergence speed of models on downstream medical tasks.




