MedNER: Covid-19 drug disease named entity recognition dataset from scientific articles
收藏官方服务:
资源简介:
This dataset contains drug and disease-named entity data for COVID-19-related texts from published papers in IOB format. This dataset has been annotated and verified by domain experts. This dataset has 26528 rows and 3 columns. The first column represents the sentence ID, the second column represents the tokens, and the third column contains the drug or disease tag. This dataset contains data for five annotated documents among 25 documents.
本数据集收录了来自已发表学术论文的新冠疫情相关文本的药物与疾病命名实体数据,采用IOB标注格式。该数据集已由领域专家完成标注与核验工作,总计包含26528条数据行与3个数据列。其中,第一列为句子ID,第二列为分词单元(Token),第三列存储药物或疾病标注标签。本数据集仅涵盖25份源文档中5份已标注文档的相关数据。
创建时间:
2024-01-23
搜集汇总
数据集介绍

背景与挑战
背景概述
MedNER是一个专注于COVID-19的药物和疾病命名实体识别数据集,数据来源于科学文章,采用IOB格式,包含26528行标注数据,由领域专家进行标注和验证。该数据集由马来西亚彭亨大学贡献,发布于2022年12月,适用于自然语言处理和生物医学应用研究。
以上内容由遇见数据集搜集并总结生成



