遇见数据集

kgnlp/meld-open-normalized

收藏
Hugging Face2026-05-12 更新2026-05-31 收录
官方服务:

资源简介:

MELD Open (Normalized) 是一个多语言命名实体识别(NER)数据集,支持超过50种语言(包括中文、英文、阿拉伯语等),规模在1000万到1亿个数据点之间。数据集包含多个子集,如AgCNER(中文农业NER,标注实体包括作物、疾病、农药等)、AgriNER(英文农业NER,标注实体包括方法、作物、疾病等)、AnatEM(英文生物医学解剖实体识别)、BC2GM(英文基因实体识别)、BC5CDR(英文化学和疾病实体识别)和BioRED(英文生物医学实体识别,涵盖基因、疾病、化学等)。数据来源于学术期刊(如PubMed、IEEE、Springer)、知识库(如中国知网、百度百科)和农业信息平台,标注由领域专家完成,采用黄金标准,部分子集标注者间一致性高(如AgCNER的Fleiss Kappa为0.966)。数据集主要用于token分类任务,适用于跨领域NER模型训练和评估。

MELD Open (Normalized) is a multilingual named entity recognition (NER) dataset supporting over 50 languages (including Chinese, English, Arabic, etc.), with a size between 10 million and 100 million data points. The dataset includes multiple subsets such as AgCNER (Chinese agricultural NER, with annotations for entities like crops, diseases, pesticides), AgriNER (English agricultural NER, with entities like methods, crops, diseases), AnatEM (English biomedical anatomical entity recognition), BC2GM (English gene entity recognition), BC5CDR (English chemical and disease entity recognition), and BioRED (English biomedical entity recognition covering genes, diseases, chemicals). Data sources include academic journals (e.g., PubMed, IEEE, Springer), knowledge bases (e.g., China National Knowledge Infrastructure, Baidu Baike), and agricultural information platforms. Annotations are gold-standard, created by domain experts, with high inter-annotator agreement in some subsets (e.g., AgCNER has a Fleiss Kappa of 0.966). The dataset is designed for token-classification tasks, suitable for training and evaluating cross-domain NER models.

提供机构:
kgnlp
二维码
社区交流群
二维码
科研交流群
商业服务