kgnlp/meld-open
收藏资源简介:
MELD Open是一个多语言、大规模的数据集,专门用于命名实体识别(NER)任务。它包含多个子集,覆盖农业和生物医学领域,支持包括中文、英文在内的多种语言。数据集结构包括训练、验证和测试分割,每个子集都有详细的元数据,如数据来源、注释协议、标签集和领域信息。例如,AgCNER子集专注于中文农业文本,包含13个实体类别;AgriNER子集针对英文农业文本,有36个实体类别;AnatEM子集处理英文生物医学解剖实体;BC2GM子集聚焦基因实体;BC5CDR子集涉及化学与疾病关系;BioRED子集用于生物医学关系提取。数据来源于学术数据库、期刊和网络资源,注释由领域专家完成,具有高一致性(如Fleiss kappa值达0.966)。数据集旨在支持跨领域NER模型的研究和应用。
MELD Open is a multilingual, large-scale dataset specifically designed for Named Entity Recognition (NER) tasks. It encompasses multiple subsets spanning agricultural and biomedical domains, and supports multiple languages including Chinese and English. The dataset structure includes training, validation, and test splits, and each subset is equipped with detailed metadata such as data sources, annotation protocols, label sets and domain information. For example, the AgCNER subset focuses on Chinese agricultural texts and contains 13 entity categories; the AgriNER subset targets English agricultural texts with 36 entity categories; the AnatEM subset handles English biomedical anatomical entities; the BC2GM subset focuses on gene entities; the BC5CDR subset involves chemical-disease relationships; and the BioRED subset is intended for biomedical relation extraction. The data is sourced from academic databases, journals and web resources, with annotations completed by domain experts, exhibiting high consistency (e.g., a Fleiss kappa score of 0.966). This dataset aims to support research and applications of cross-domain NER models.



